{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/an-end-to-end-textspotter-with-explicit","title":"An end-to-end TextSpotter with Explicit Alignment and Attention","arxiv_id":"1803.03474","date":"2018-03-09","proceeding":"CVPR 2018 6","authors":["Tong He","Zhi Tian","Weilin Huang","Chunhua Shen","Yu Qiao","Changming Sun"],"abstract":"Text detection and recognition in natural images have long been considered as\ntwo separate tasks that are processed sequentially. Training of two tasks in a\nunified framework is non-trivial due to significant dif- ferences in\noptimisation difficulties. In this work, we present a conceptually simple yet\nefficient framework that simultaneously processes the two tasks in one shot.\nOur main contributions are three-fold: 1) we propose a novel text-alignment\nlayer that allows it to precisely compute convolutional features of a text\ninstance in ar- bitrary orientation, which is the key to boost the per-\nformance; 2) a character attention mechanism is introduced by using character\nspatial information as explicit supervision, leading to large improvements in\nrecognition; 3) two technologies, together with a new RNN branch for word\nrecognition, are integrated seamlessly into a single model which is end-to-end\ntrainable. This allows the two tasks to work collaboratively by shar- ing\nconvolutional features, which is critical to identify challenging text\ninstances. Our model achieves impressive results in end-to-end recognition on\nthe ICDAR2015 dataset, significantly advancing most recent results, with\nimprovements of F-measure from (0.54, 0.51, 0.47) to (0.82, 0.77, 0.63), by\nusing a strong, weak and generic lexicon respectively. Thanks to joint\ntraining, our method can also serve as a good detec- tor by achieving a new\nstate-of-the-art detection performance on two datasets.","url_abs":"http://arxiv.org/abs/1803.03474v3","url_pdf":"http://arxiv.org/pdf/1803.03474v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"an-end-to-end-textspotter-with-explicit","repo_url":"https://github.com/tonghe90/textspotter","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"an-end-to-end-textspotter-with-explicit","repo_url":"https://github.com/curbmap/curbmap-ml","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"text-detection","task_name":"Text Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1803.03474","atlas_url":"https://app.syntology.ai/?focus=1803.03474","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}