{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/transfer-learning-for-ocropus-model-training","title":"Transfer Learning for OCRopus Model Training on Early Printed Books","arxiv_id":"1712.05586","date":"2017-12-15","proceeding":null,"authors":["Christian Reul","Christoph Wick","Uwe Springmann","Frank Puppe"],"abstract":"A method is presented that significantly reduces the character error rates\nfor OCR text obtained from OCRopus models trained on early printed books when\nonly small amounts of diplomatic transcriptions are available. This is achieved\nby building from already existing models during training instead of starting\nfrom scratch. To overcome the discrepancies between the set of characters of\nthe pretrained model and the additional ground truth the OCRopus code is\nadapted to allow for alphabet expansion or reduction. The character set is now\ncapable of flexibly adding and deleting characters from the pretrained alphabet\nwhen an existing model is loaded. For our experiments we use a self-trained\nmixed model on early Latin prints and the two standard OCRopus models on modern\nEnglish and German Fraktur texts. The evaluation on seven early printed books\nshowed that training from the Latin mixed model reduces the average amount of\nerrors by 43% and 26%, respectively compared to training from scratch with 60\nand 150 lines of ground truth, respectively. Furthermore, it is shown that even\nbuilding from mixed models trained on data unrelated to the newly added\ntraining and test data can lead to significantly improved recognition results.","url_abs":"http://arxiv.org/abs/1712.05586v2","url_pdf":"http://arxiv.org/pdf/1712.05586v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"transfer-learning-for-ocropus-model-training","repo_url":"https://github.com/chreul/OCR_Testdata_EarlyPrintedBooks","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"optical-character-recognition","task_name":"Optical Character Recognition (OCR)"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}