{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/scene-text-recognition-from-two-dimensional","title":"Scene Text Recognition from Two-Dimensional Perspective","arxiv_id":"1809.06508","date":"2018-09-18","proceeding":null,"authors":["Minghui Liao","Jian Zhang","Zhaoyi Wan","Fengming Xie","Jiajun Liang","Pengyuan Lyu","Cong Yao","Xiang Bai"],"abstract":"Inspired by speech recognition, recent state-of-the-art algorithms mostly\nconsider scene text recognition as a sequence prediction problem. Though\nachieving excellent performance, these methods usually neglect an important\nfact that text in images are actually distributed in two-dimensional space. It\nis a nature quite different from that of speech, which is essentially a\none-dimensional signal. In principle, directly compressing features of text\ninto a one-dimensional form may lose useful information and introduce extra\nnoise. In this paper, we approach scene text recognition from a two-dimensional\nperspective. A simple yet effective model, called Character Attention Fully\nConvolutional Network (CA-FCN), is devised for recognizing the text of\narbitrary shapes. Scene text recognition is realized with a semantic\nsegmentation network, where an attention mechanism for characters is adopted.\nCombined with a word formation module, CA-FCN can simultaneously recognize the\nscript and predict the position of each character. Experiments demonstrate that\nthe proposed algorithm outperforms previous methods on both regular and\nirregular text datasets. Moreover, it is proven to be more robust to imprecise\nlocalizations in the text detection phase, which are very common in practice.","url_abs":"http://arxiv.org/abs/1809.06508v2","url_pdf":"http://arxiv.org/pdf/1809.06508v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"scene-text-recognition","task_name":"Scene Text Recognition"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"text-detection","task_name":"Text Detection"},{"task_slug":"two","task_name":"Vocal Bursts Valence Prediction"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/scene-text-recognition-on-icdar2013","task":"Scene Text Recognition","dataset":"ICDAR2013","model":"CA-FCN","rank_in_archive_order":33,"of":38,"metrics":{"Accuracy":"91.5"},"uses_additional_data":false},{"leaderboard":"/sota/scene-text-recognition-on-svt","task":"Scene Text Recognition","dataset":"SVT","model":"CA-FCN","rank_in_archive_order":32,"of":37,"metrics":{"Accuracy":"86.4"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1809.06508","atlas_url":"https://app.syntology.ai/?focus=1809.06508","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}