{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-multi-object-rectified-attention-network","title":"A Multi-Object Rectified Attention Network for Scene Text Recognition","arxiv_id":"1901.03003","date":"2019-01-10","proceeding":null,"authors":["Canjie Luo","Lianwen Jin","Zenghui Sun"],"abstract":"Irregular text is widely used. However, it is considerably difficult to\nrecognize because of its various shapes and distorted patterns. In this paper,\nwe thus propose a multi-object rectified attention network (MORAN) for general\nscene text recognition. The MORAN consists of a multi-object rectification\nnetwork and an attention-based sequence recognition network. The multi-object\nrectification network is designed for rectifying images that contain irregular\ntext. It decreases the difficulty of recognition and enables the\nattention-based sequence recognition network to more easily read irregular\ntext. It is trained in a weak supervision way, thus requiring only images and\ncorresponding text labels. The attention-based sequence recognition network\nfocuses on target characters and sequentially outputs the predictions.\nMoreover, to improve the sensitivity of the attention-based sequence\nrecognition network, a fractional pickup method is proposed for an\nattention-based decoder in the training phase. With the rectification\nmechanism, the MORAN can read both regular and irregular scene text. Extensive\nexperiments on various benchmarks are conducted, which show that the MORAN\nachieves state-of-the-art performance. The source code is available.","url_abs":"http://arxiv.org/abs/1901.03003v1","url_pdf":"http://arxiv.org/pdf/1901.03003v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-multi-object-rectified-attention-network","repo_url":"https://github.com/Canjie-Luo/MORAN_v2","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"a-multi-object-rectified-attention-network","repo_url":"https://github.com/ModelBunker/MORAN-PyTorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"a-multi-object-rectified-attention-network","repo_url":"https://github.com/agiletechvn/moran_v2_text_recognition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"a-multi-object-rectified-attention-network","repo_url":"https://github.com/dipu-bd/craft-moran-ocr","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"a-multi-object-rectified-attention-network","repo_url":"https://github.com/jeasung-pf/MORAN_v2","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"a-multi-object-rectified-attention-network","repo_url":"https://github.com/lzmisscc/emoran","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"a-multi-object-rectified-attention-network","repo_url":"https://github.com/topdu/openocr","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"object","task_name":"Object"},{"task_slug":"optical-character-recognition","task_name":"Optical Character Recognition (OCR)"},{"task_slug":"scene-text-recognition","task_name":"Scene Text Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/optical-character-recognition-on-benchmarking","task":"Optical Character Recognition (OCR)","dataset":"Benchmarking Chinese Text Recognition: Datasets, Baselines, and an Empirical Study","model":"MORAN","rank_in_archive_order":6,"of":7,"metrics":{"Accuracy (%)":"64.3"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}