{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/tmixt-a-process-flow-for-transcribing-mixed","title":"TMIXT: A process flow for Transcribing MIXed handwritten and machine-printed Text","arxiv_id":"1904.12387","date":"2019-04-28","proceeding":null,"authors":["Fady Medhat","Mahnaz Mohammadi","Sardar Jaf","Chris G. Willcocks","Toby P. Breckon","Peter Matthews","Andrew Stephen McGough","Georgios Theodoropoulos","Boguslaw Obara"],"abstract":"Handling large corpuses of documents is of significant importance in many\nfields, no more so than in the areas of crime investigation and defence, where\nan organisation may be presented with a large volume of scanned documents which\nneed to be processed in a finite time. However, this problem is exacerbated\nboth by the volume, in terms of scanned documents and the complexity of the\npages, which need to be processed. Often containing many different elements,\nwhich each need to be processed and understood. Text recognition, which is a\nprimary task of this process, is usually dependent upon the type of text, being\neither handwritten or machine-printed. Accordingly, the recognition involves\nprior classification of the text category, before deciding on the recognition\nmethod to be applied. This poses a more challenging task if a document contains\nboth handwritten and machine-printed text. In this work, we present a generic\nprocess flow for text recognition in scanned documents containing mixed\nhandwritten and machine-printed text without the need to classify text in\nadvance. We realize the proposed process flow using several open-source image\nprocessing and text recognition packages1. The evaluation is performed using a\nspecially developed variant, presented in this work, of the IAM handwriting\ndatabase, where we achieve an average transcription accuracy of nearly 80% for\npages containing both printed and handwritten text.","url_abs":"http://arxiv.org/abs/1904.12387v1","url_pdf":"http://arxiv.org/pdf/1904.12387v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"tmixt-a-process-flow-for-transcribing-mixed","repo_url":"https://github.com/fadymedhat/TMIXT","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}