{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/docemul-a-toolkit-to-generate-structured","title":"DocEmul: a Toolkit to Generate Structured Historical Documents","arxiv_id":"1710.03474","date":"2017-10-10","proceeding":null,"authors":["Samuele Capobianco","Simone Marinai"],"abstract":"We propose a toolkit to generate structured synthetic documents emulating the\nactual document production process. Synthetic documents can be used to train\nsystems to perform document analysis tasks. In our case we address the record\ncounting task on handwritten structured collections containing a limited number\nof examples. Using the DocEmul toolkit we can generate a larger dataset to\ntrain a deep architecture to predict the number of records for each page. The\ntoolkit is able to generate synthetic collections and also perform data\naugmentation to create a larger trainable dataset. It includes one method to\nextract the page background from real pages which can be used as a substrate\nwhere records can be written on the basis of variable structures and using\ncursive fonts. Moreover, it is possible to extend the synthetic collection by\nadding random noise, page rotations, and other visual variations. We performed\nsome experiments on two different handwritten collections using the toolkit to\ngenerate synthetic data to train a Convolutional Neural Network able to count\nthe number of records in the real collections.","url_abs":"http://arxiv.org/abs/1710.03474v1","url_pdf":"http://arxiv.org/pdf/1710.03474v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"docemul-a-toolkit-to-generate-structured","repo_url":"https://github.com/scstech85/DocEmul","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"data-augmentation","task_name":"Data Augmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1710.03474","atlas_url":"https://app.syntology.ai/?focus=1710.03474","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}