{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/single-shot-text-detector-with-regional","title":"Single Shot Text Detector with Regional Attention","arxiv_id":"1709.00138","date":"2017-09-01","proceeding":"ICCV 2017 10","authors":["Pan He","Weilin Huang","Tong He","Qile Zhu","Yu Qiao","Xiaolin Li"],"abstract":"We present a novel single-shot text detector that directly outputs word-level\nbounding boxes in a natural image. We propose an attention mechanism which\nroughly identifies text regions via an automatically learned attentional map.\nThis substantially suppresses background interference in the convolutional\nfeatures, which is the key to producing accurate inference of words,\nparticularly at extremely small sizes. This results in a single model that\nessentially works in a coarse-to-fine manner. It departs from recent FCN- based\ntext detectors which cascade multiple FCN models to achieve an accurate\nprediction. Furthermore, we develop a hierarchical inception module which\nefficiently aggregates multi-scale inception features. This enhances local\ndetails, and also encodes strong context information, allow- ing the detector\nto work reliably on multi-scale and multi- orientation text with single-scale\nimages. Our text detector achieves an F-measure of 77% on the ICDAR 2015 bench-\nmark, advancing the state-of-the-art results in [18, 28]. Demo is available at:\nhttp://sstd.whuang.org/.","url_abs":"http://arxiv.org/abs/1709.00138v1","url_pdf":"http://arxiv.org/pdf/1709.00138v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"single-shot-text-detector-with-regional","repo_url":"https://github.com/BestSonny/SSTD","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"scene-text-detection","task_name":"Scene Text Detection"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"fcn","method_name":"FCN"},{"method_slug":"inception-module","method_name":"Inception Module"},{"method_slug":"max-pooling","method_name":"Max Pooling"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/scene-text-detection-on-coco-text","task":"Scene Text Detection","dataset":"COCO-Text","model":"SSTD","rank_in_archive_order":4,"of":6,"metrics":{"F-Measure":"37","Precision":"46","Recall":"31"},"uses_additional_data":false},{"leaderboard":"/sota/scene-text-detection-on-icdar-2013","task":"Scene Text Detection","dataset":"ICDAR 2013","model":"SSTD","rank_in_archive_order":10,"of":16,"metrics":{"F-Measure":"87%","Precision":"88","Recall":"86"},"uses_additional_data":false},{"leaderboard":"/sota/scene-text-detection-on-icdar-2015","task":"Scene Text Detection","dataset":"ICDAR 2015","model":"EAST + PVANET2x RBOX (multi-scale)","rank_in_archive_order":36,"of":43,"metrics":{"F-Measure":"80.7","Precision":"83.3","Recall":"78.3"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1709.00138","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}