{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/case-study-of-a-highly-automated-layout","title":"Case Study of a highly automated Layout Analysis and OCR of an incunabulum: 'Der Heiligen Leben' (1488)","arxiv_id":"1701.07395","date":"2017-01-20","proceeding":null,"authors":["Christian Reul","Marco Dittrich","Martin Gruner"],"abstract":"This paper provides the first thorough documentation of a high quality\ndigitization process applied to an early printed book from the incunabulum\nperiod (1450-1500). The entire OCR related workflow including preprocessing,\nlayout analysis and text recognition is illustrated in detail using the example\nof 'Der Heiligen Leben', printed in Nuremberg in 1488. For each step the\nrequired time expenditure was recorded. The character recognition yielded\nexcellent results both on character (97.57%) and word (92.19%) level.\nFurthermore, a comparison of a highly automated (LAREX) and a manual (Aletheia)\nmethod for layout analysis was performed. By considerably automating the\nsegmentation the required human effort was reduced significantly from over 100\nhours to less than six hours, resulting in only a slight drop in OCR accuracy.\nRealistic estimates for the human effort necessary for full text extraction\nfrom incunabula can be derived from this study. The printed pages of the\ncomplete work together with the OCR result is available online ready to be\ninspected and downloaded.","url_abs":"http://arxiv.org/abs/1701.07395v1","url_pdf":"http://arxiv.org/pdf/1701.07395v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"case-study-of-a-highly-automated-layout","repo_url":"https://github.com/OCR4all/LAREX","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"case-study-of-a-highly-automated-layout","repo_url":"https://github.com/chreul/LAREX","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"optical-character-recognition","task_name":"Optical Character Recognition (OCR)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}