Skip to content

Detail of publication

Citation

Ivan Gruber and Pavel Ircing and Petr Neduchal and Marek Hrúz and Miroslav Hlaváč and Zbyněk Zajíc and Jan Švec and Martin Bulín : An Automated Pipeline for Robust Image Processing and Optical Character Recognition of Historical Documents . SPECOM: International Conference on Speech and Computer, Lecture Notes in Computer Science , vol. 12335, p. 166-175, Springer, Cham, 2020.

Abstract

In this paper we propose a pipeline for processing of scanned historical documents into the electronic text form that could then be indexed and stored in a database. The nature of the documents presents a substantial challenge for standard automated techniques – not only there is a mix of typewritten and handwritten documents of varying quality but the scanned pages often contain multiple documents at once. Moreover, the language of the texts alternates mostly between Russian and Ukrainian but other languages also occur. The paper focuses mainly on segmentation, document type classification, and image preprocessing of the scanned documents; the output of those methods is then passed to the off-the-shelf OCR software and a baseline performance is evaluated on a simplified OCR task.

Detail of publication

Title: An Automated Pipeline for Robust Image Processing and Optical Character Recognition of Historical Documents
Author: Ivan Gruber ; Pavel Ircing ; Petr Neduchal ; Marek Hrúz ; Miroslav Hlaváč ; Zbyněk Zajíc ; Jan Švec ; Martin Bulín
Language: English
Date of publication: 29 Sep 2020
Year: 2020
Type of publication: Papers in proceedings of reviewed conferences
Title of journal or book: SPECOM: International Conference on Speech and Computer
Series: Lecture Notes in Computer Science
Číslo vydání: 12335
Page: 166 - 175
DOI: https://doi.org/10.1007/978-3-030-60276-5_17
ISBN: 978-3-030-60276-5
ISSN: 1611-3349
Publisher: Springer, Cham
/ 2021-01-16 09:52:19 /

BibTeX

@INPROCEEDINGS{IvanGruber_2020_AnAutomatedPipeline,
 author = {Ivan Gruber and Pavel Ircing and Petr Neduchal and Marek Hr\'{u}z and Miroslav Hlav\'{a}\v{c} and Zbyn\v{e}k Zaj\'{i}c and Jan \v{S}vec and Martin Bul\'{i}n},
 title = {An Automated Pipeline for Robust Image Processing and Optical Character Recognition of Historical Documents},
 year = {2020},
 publisher = {Springer, Cham},
 journal = {SPECOM: International Conference on Speech and Computer},
 volume = {12335},
 pages = {166-175},
 series = {Lecture Notes in Computer Science },
 ISBN = {978-3-030-60276-5},
 ISSN = {1611-3349},
 doi = {https://doi.org/10.1007/978-3-030-60276-5_17},
 url = {http://www.kky.zcu.cz/en/publications/IvanGruber_2020_AnAutomatedPipeline},
}