{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/scene-text-detection-via-holistic-multi","title":"Scene Text Detection via Holistic, Multi-Channel Prediction","arxiv_id":"1606.09002","date":"2016-06-29","proceeding":null,"authors":["Cong Yao","Xiang Bai","Nong Sang","Xinyu Zhou","Shuchang Zhou","Zhimin Cao"],"abstract":"Recently, scene text detection has become an active research topic in\ncomputer vision and document analysis, because of its great importance and\nsignificant challenge. However, vast majority of the existing methods detect\ntext within local regions, typically through extracting character, word or line\nlevel candidates followed by candidate aggregation and false positive\nelimination, which potentially exclude the effect of wide-scope and long-range\ncontextual cues in the scene. To take full advantage of the rich information\navailable in the whole natural image, we propose to localize text in a holistic\nmanner, by casting scene text detection as a semantic segmentation problem. The\nproposed algorithm directly runs on full images and produces global, pixel-wise\nprediction maps, in which detections are subsequently formed. To better make\nuse of the properties of text, three types of information regarding text\nregion, individual characters and their relationship are estimated, with a\nsingle Fully Convolutional Network (FCN) model. With such predictions of text\nproperties, the proposed algorithm can simultaneously handle horizontal,\nmulti-oriented and curved text in real-world natural images. The experiments on\nstandard benchmarks, including ICDAR 2013, ICDAR 2015 and MSRA-TD500,\ndemonstrate that the proposed algorithm substantially outperforms previous\nstate-of-the-art approaches. Moreover, we report the first baseline result on\nthe recently-released, large-scale dataset COCO-Text.","url_abs":"http://arxiv.org/abs/1606.09002v2","url_pdf":"http://arxiv.org/pdf/1606.09002v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"prediction","task_name":"Prediction"},{"task_slug":"scene-text-detection","task_name":"Scene Text Detection"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"text-detection","task_name":"Text Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/scene-text-detection-on-coco-text","task":"Scene Text Detection","dataset":"COCO-Text","model":"Yao et al.","rank_in_archive_order":6,"of":6,"metrics":{"F-Measure":"33.31","Precision":"43.23","Recall":"27.1"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1606.09002","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}