{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/wordsup-exploiting-word-annotations-for","title":"WordSup: Exploiting Word Annotations for Character based Text Detection","arxiv_id":"1708.06720","date":"2017-08-22","proceeding":"ICCV 2017 10","authors":["Han Hu","Chengquan Zhang","Yuxuan Luo","Yuzhuo Wang","Junyu Han","Errui Ding"],"abstract":"Imagery texts are usually organized as a hierarchy of several visual\nelements, i.e. characters, words, text lines and text blocks. Among these\nelements, character is the most basic one for various languages such as\nWestern, Chinese, Japanese, mathematical expression and etc. It is natural and\nconvenient to construct a common text detection engine based on character\ndetectors. However, training character detectors requires a vast of location\nannotated characters, which are expensive to obtain. Actually, the existing\nreal text datasets are mostly annotated in word or line level. To remedy this\ndilemma, we propose a weakly supervised framework that can utilize word\nannotations, either in tight quadrangles or the more loose bounding boxes, for\ncharacter detector training. When applied in scene text detection, we are thus\nable to train a robust character detector by exploiting word annotations in the\nrich large-scale real scene text datasets, e.g. ICDAR15 and COCO-text. The\ncharacter detector acts as a key role in the pipeline of our text detection\nengine. It achieves the state-of-the-art performance on several challenging\nscene text detection benchmarks. We also demonstrate the flexibility of our\npipeline by various scenarios, including deformed text detection and math\nexpression recognition.","url_abs":"http://arxiv.org/abs/1708.06720v1","url_pdf":"http://arxiv.org/pdf/1708.06720v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"math","task_name":"Math"},{"task_slug":"scene-text-detection","task_name":"Scene Text Detection"},{"task_slug":"text-detection","task_name":"Text Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/scene-text-detection-on-coco-text","task":"Scene Text Detection","dataset":"COCO-Text","model":"WordSup (VGG16-synth-coco)","rank_in_archive_order":5,"of":6,"metrics":{"F-Measure":"36.8","Precision":"45.2","Recall":"30.9"},"uses_additional_data":false},{"leaderboard":"/sota/scene-text-detection-on-icdar-2013","task":"Scene Text Detection","dataset":"ICDAR 2013","model":"WordSup (VGG16-synth-icdar)","rank_in_archive_order":4,"of":16,"metrics":{"F-Measure":"90.34%","Precision":"93.34","Recall":"87.53"},"uses_additional_data":false},{"leaderboard":"/sota/scene-text-detection-on-icdar-2015","task":"Scene Text Detection","dataset":"ICDAR 2015","model":"SSTD","rank_in_archive_order":39,"of":43,"metrics":{"F-Measure":"77","Precision":"80","Recall":"73"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}