{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/aon-towards-arbitrarily-oriented-text","title":"AON: Towards Arbitrarily-Oriented Text Recognition","arxiv_id":"1711.04226","date":"2017-11-12","proceeding":"CVPR 2018 6","authors":["Zhanzhan Cheng","Yangliu Xu","Fan Bai","Yi Niu","ShiLiang Pu","Shuigeng Zhou"],"abstract":"Recognizing text from natural images is a hot research topic in computer\nvision due to its various applications. Despite the enduring research of\nseveral decades on optical character recognition (OCR), recognizing texts from\nnatural images is still a challenging task. This is because scene texts are\noften in irregular (e.g. curved, arbitrarily-oriented or seriously distorted)\narrangements, which have not yet been well addressed in the literature.\nExisting methods on text recognition mainly work with regular (horizontal and\nfrontal) texts and cannot be trivially generalized to handle irregular texts.\nIn this paper, we develop the arbitrary orientation network (AON) to directly\ncapture the deep features of irregular texts, which are combined into an\nattention-based decoder to generate character sequence. The whole network can\nbe trained end-to-end by using only images and word-level annotations.\nExtensive experiments on various benchmarks, including the CUTE80,\nSVT-Perspective, IIIT5k, SVT and ICDAR datasets, show that the proposed\nAON-based method achieves the-state-of-the-art performance in irregular\ndatasets, and is comparable to major existing methods in regular datasets.","url_abs":"http://arxiv.org/abs/1711.04226v2","url_pdf":"http://arxiv.org/pdf/1711.04226v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"aon-towards-arbitrarily-oriented-text","repo_url":"https://github.com/huizhang0110/AON","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"optical-character-recognition","task_name":"Optical Character Recognition"},{"task_slug":"optical-character-recognition","task_name":"Optical Character Recognition (OCR)"},{"task_slug":"scene-text-recognition","task_name":"Scene Text Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/scene-text-recognition-on-icdar-2003","task":"Scene Text Recognition","dataset":"ICDAR 2003","model":"AON","rank_in_archive_order":9,"of":12,"metrics":{"Accuracy":"91.5"},"uses_additional_data":false},{"leaderboard":"/sota/scene-text-recognition-on-icdar2015","task":"Scene Text Recognition","dataset":"ICDAR2015","model":"AON","rank_in_archive_order":24,"of":27,"metrics":{"Accuracy":"73.0"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}