{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/synthetic-data-and-artificial-neural-networks","title":"Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition","arxiv_id":"1406.2227","date":"2014-06-09","proceeding":null,"authors":["Max Jaderberg","Karen Simonyan","Andrea Vedaldi","Andrew Zisserman"],"abstract":"In this work we present a framework for the recognition of natural scene\ntext. Our framework does not require any human-labelled data, and performs word\nrecognition on the whole image holistically, departing from the character based\nrecognition systems of the past. The deep neural network models at the centre\nof this framework are trained solely on data produced by a synthetic text\ngeneration engine -- synthetic data that is highly realistic and sufficient to\nreplace real data, giving us infinite amounts of training data. This excess of\ndata exposes new possibilities for word recognition models, and here we\nconsider three models, each one \"reading\" words in a different way: via 90k-way\ndictionary encoding, character sequence encoding, and bag-of-N-grams encoding.\nIn the scenarios of language based and completely unconstrained text\nrecognition we greatly improve upon state-of-the-art performance on standard\ndatasets, using our fast, simple machinery and requiring zero data-acquisition\ncosts.","url_abs":"http://arxiv.org/abs/1406.2227v4","url_pdf":"http://arxiv.org/pdf/1406.2227v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"synthetic-data-and-artificial-neural-networks","repo_url":"https://github.com/bupt-ai-cz/Meta-SelfLearning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"scene-text-recognition","task_name":"Scene Text Recognition"},{"task_slug":"text-generation","task_name":"Text Generation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/scene-text-recognition-on-icdar2013","task":"Scene Text Recognition","dataset":"ICDAR2013","model":"CHAR","rank_in_archive_order":38,"of":38,"metrics":{"Accuracy":"79.5"},"uses_additional_data":false},{"leaderboard":"/sota/scene-text-recognition-on-svt","task":"Scene Text Recognition","dataset":"SVT","model":"CHAR","rank_in_archive_order":37,"of":37,"metrics":{"Accuracy":"68.0"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1406.2227","atlas_url":"https://app.syntology.ai/?focus=1406.2227","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}