{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/accurate-data-efficient-unconstrained-text","title":"Accurate, Data-Efficient, Unconstrained Text Recognition with Convolutional Neural Networks","arxiv_id":"1812.11894","date":"2018-12-31","proceeding":null,"authors":["Mohamed Yousef","Khaled F. Hussain","Usama S. Mohammed"],"abstract":"Unconstrained text recognition is an important computer vision task,\nfeaturing a wide variety of different sub-tasks, each with its own set of\nchallenges. One of the biggest promises of deep neural networks has been the\nconvergence and automation of feature extractors from input raw signals,\nallowing for the highest possible performance with minimum required domain\nknowledge. To this end, we propose a data-efficient, end-to-end neural network\nmodel for generic, unconstrained text recognition. In our proposed architecture\nwe strive for simplicity and efficiency without sacrificing recognition\naccuracy. Our proposed architecture is a fully convolutional network without\nany recurrent connections trained with the CTC loss function. Thus it operates\non arbitrary input sizes and produces strings of arbitrary length in a very\nefficient and parallelizable manner. We show the generality and superiority of\nour proposed text recognition architecture by achieving state of the art\nresults on seven public benchmark datasets, covering a wide spectrum of text\nrecognition tasks, namely: Handwriting Recognition, CAPTCHA recognition, OCR,\nLicense Plate Recognition, and Scene Text Recognition. Our proposed\narchitecture has won the ICFHR2018 Competition on Automated Text Recognition on\na READ Dataset.","url_abs":"http://arxiv.org/abs/1812.11894v1","url_pdf":"http://arxiv.org/pdf/1812.11894v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"accurate-data-efficient-unconstrained-text","repo_url":"https://github.com/IntuitionMachines/OrigamiNet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"handwriting-recognition","task_name":"Handwriting Recognition"},{"task_slug":"license-plate-recognition","task_name":"License Plate Recognition"},{"task_slug":"optical-character-recognition","task_name":"Optical Character Recognition (OCR)"},{"task_slug":"scene-text-recognition","task_name":"Scene Text Recognition"}],"methods":[{"method_slug":"ctc-loss","method_name":"CTC Loss"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}