{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/robust-scene-text-recognition-with-automatic","title":"Robust Scene Text Recognition with Automatic Rectification","arxiv_id":"1603.03915","date":"2016-03-12","proceeding":"CVPR 2016 6","authors":["Baoguang Shi","Xinggang Wang","Pengyuan Lyu","Cong Yao","Xiang Bai"],"abstract":"Recognizing text in natural images is a challenging task with many unsolved\nproblems. Different from those in documents, words in natural images often\npossess irregular shapes, which are caused by perspective distortion, curved\ncharacter placement, etc. We propose RARE (Robust text recognizer with\nAutomatic REctification), a recognition model that is robust to irregular text.\nRARE is a specially-designed deep neural network, which consists of a Spatial\nTransformer Network (STN) and a Sequence Recognition Network (SRN). In testing,\nan image is firstly rectified via a predicted Thin-Plate-Spline (TPS)\ntransformation, into a more \"readable\" image for the following SRN, which\nrecognizes text through a sequence recognition approach. We show that the model\nis able to recognize several types of irregular text, including perspective\ntext and curved text. RARE is end-to-end trainable, requiring only images and\nassociated text labels, making it convenient to train and deploy the model in\npractical systems. State-of-the-art or highly-competitive performance achieved\non several benchmarks well demonstrates the effectiveness of the proposed\nmodel.","url_abs":"http://arxiv.org/abs/1603.03915v2","url_pdf":"http://arxiv.org/pdf/1603.03915v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"robust-scene-text-recognition-with-automatic","repo_url":"https://github.com/Media-Smart/vedastr","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"robust-scene-text-recognition-with-automatic","repo_url":"https://github.com/PaddlePaddle/PaddleOCR","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"paddle","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"robust-scene-text-recognition-with-automatic","repo_url":"https://github.com/WarBean/tps_stn_pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"robust-scene-text-recognition-with-automatic","repo_url":"https://github.com/iwyoo/tf_thinplatespline","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"robust-scene-text-recognition-with-automatic","repo_url":"https://github.com/mindspore-lab/mindocr","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":{"status":"gone","observed_at":"2026-09-17","how":"tree_404+repo_404"}}],"tasks":[{"task_slug":"optical-character-recognition","task_name":"Optical Character Recognition (OCR)"},{"task_slug":"scene-text-detection","task_name":"Scene Text Detection"},{"task_slug":"scene-text-recognition","task_name":"Scene Text Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/scene-text-recognition-on-icdar-2003","task":"Scene Text Recognition","dataset":"ICDAR 2003","model":"RARE","rank_in_archive_order":10,"of":12,"metrics":{"Accuracy":"90.1"},"uses_additional_data":false},{"leaderboard":"/sota/scene-text-recognition-on-icdar2013","task":"Scene Text Recognition","dataset":"ICDAR2013","model":"RARE","rank_in_archive_order":36,"of":38,"metrics":{"Accuracy":"88.6"},"uses_additional_data":false},{"leaderboard":"/sota/scene-text-recognition-on-svt","task":"Scene Text Recognition","dataset":"SVT","model":"RARE","rank_in_archive_order":35,"of":37,"metrics":{"Accuracy":"81.9"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1603.03915","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}