{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/star-net-a-spatial-attention-residue-network","title":"Star-net: A spatial attention residue network for scene text recognition.","arxiv_id":null,"date":"2016-09-20","proceeding":"The British Machine Vision Conference,2016 2016 9","authors":["W. Liu","C. Chen","K.-Y. K. Wong","Z. Su","and J. Han."],"abstract":"In this paper, we present a novel SpaTial Attention Residue Network (STAR-Net)\r\nfor recognising scene texts. Our STAR-Net is equipped with a spatial attention mechanism which employs a spatial transformer to remove the distortions of texts in natural\r\nimages. This allows the subsequent feature extractor to focus on the rectified text region without being sidetracked by the distortions. Our STAR-Net also exploits residue\r\nconvolutional blocks to build a very deep feature extractor, which is essential to the successful extraction of discriminative text features for this fine grained recognition task.\r\nCombining the spatial attention mechanism with the residue convolutional blocks, our\r\nSTAR-Net is the deepest end-to-end trainable neural network for scene text recognition.\r\nExperiments have been conducted on five public benchmark datasets. Experimental results show that our STAR-Net can achieve a performance comparable to state-of-the-art\r\nmethods for scene texts with little distortions, and outperform these methods for scene\r\ntexts with considerable distortions.","url_abs":"http://www.bmva.org/bmvc/2016/papers/paper043/paper043.pdf","url_pdf":"http://www.bmva.org/bmvc/2016/papers/paper043/paper043.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"star-net-a-spatial-attention-residue-network","repo_url":"https://github.com/PaddlePaddle/PaddleOCR","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"paddle","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"optical-character-recognition","task_name":"Optical Character Recognition (OCR)"},{"task_slug":"scene-text-recognition","task_name":"Scene Text Recognition"}],"methods":[{"method_slug":"spatial-transformer","method_name":"Spatial Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/scene-text-recognition-on-icdar-2003","task":"Scene Text Recognition","dataset":"ICDAR 2003","model":"STAR-Net","rank_in_archive_order":11,"of":12,"metrics":{"Accuracy":"89.9"},"uses_additional_data":false},{"leaderboard":"/sota/scene-text-recognition-on-icdar2013","task":"Scene Text Recognition","dataset":"ICDAR2013","model":"STAR-Net","rank_in_archive_order":35,"of":38,"metrics":{"Accuracy":"89.1"},"uses_additional_data":false},{"leaderboard":"/sota/scene-text-recognition-on-svt","task":"Scene Text Recognition","dataset":"SVT","model":"STAR-Net","rank_in_archive_order":34,"of":37,"metrics":{"Accuracy":"83.6"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}