Papers › WordSup: Exploiting Word Annotations for Character based Text Detection

WordSup: Exploiting Word Annotations for Character based Text Detection

22 Aug 2017ICCV 2017 10arXiv:1708.06720archive 2025-07-28

Han Hu, Chengquan Zhang, Yuxuan Luo, Yuzhuo Wang, Junyu Han, Errui Ding

Imagery texts are usually organized as a hierarchy of several visual elements, i.e. characters, words, text lines and text blocks. Among these elements, character is the most basic one for various languages such as Western, Chinese, Japanese, mathematical expression and etc. It is natural and convenient to construct a common text detection engine based on character detectors. However, training character detectors requires a vast of location annotated characters, which are expensive to obtain. Actually, the existing real text datasets are mostly annotated in word or line level. To remedy this dilemma, we propose a weakly supervised framework that can utilize word annotations, either in tight quadrangles or the more loose bounding boxes, for character detector training. When applied in scene text detection, we are thus able to train a robust character detector by exploiting word annotations in the rich large-scale real scene text datasets, e.g. ICDAR15 and COCO-text. The character detector acts as a key role in the pipeline of our text detection engine. It achieves the state-of-the-art performance on several challenging scene text detection benchmarks. We also demonstrate the flexibility of our pipeline by various scenarios, including deformed text detection and math expression recognition.

PaperPDFConference PDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

MathScene Text DetectionText Detection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Scene Text Detection COCO-Text WordSup (VGG16-synth-coco) F-Measure 36.8 #5 of 6 Archive leaderboard report
Scene Text Detection COCO-Text WordSup (VGG16-synth-coco) Precision 45.2 #5 of 6 Archive leaderboard report
Scene Text Detection COCO-Text WordSup (VGG16-synth-coco) Recall 30.9 #5 of 6 Archive leaderboard report
Scene Text Detection ICDAR 2013 WordSup (VGG16-synth-icdar) F-Measure 90.34% #4 of 16 Archive leaderboard report
Scene Text Detection ICDAR 2013 WordSup (VGG16-synth-icdar) Precision 93.34 #4 of 16 Archive leaderboard report
Scene Text Detection ICDAR 2013 WordSup (VGG16-synth-icdar) Recall 87.53 #4 of 16 Archive leaderboard report
Scene Text Detection ICDAR 2015 SSTD F-Measure 77 #39 of 43 Archive leaderboard report
Scene Text Detection ICDAR 2015 SSTD Precision 80 #39 of 43 Archive leaderboard report
Scene Text Detection ICDAR 2015 SSTD Recall 73 #39 of 43 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections