Papers › TextFuseNet: Scene Text Detection with Richer Fused Features

TextFuseNet: Scene Text Detection with Richer Fused Features

17 May 2020archive 2025-07-28

Jian Ye, Zhe Chen, Juhua Liu, Bo Du

Arbitrary shape text detection in natural scenes is an extremely challenging task. Unlike existing text detection approaches that only perceive texts based on limited feature representations, we propose a novel framework, namely TextFuseNet, to exploit the use of richer features fused for text detection. More specifically, we propose to perceive texts from three levels of feature representations, i.e., character-, word- and global-level, and then introduce a novel text representation fusion technique to help achieve robust arbitrary text detection. The multi-level feature representation can adequately describe texts by dissecting them into individual characters while still maintaining their general semantics. TextFuseNet then collects and merges the texts’ features from different levels using a multi-path fusion architecture which can effectively align and fuse different representations. In practice, our proposed TextFuseNet can learn a more adequate description of arbitrary shapes texts, suppressing false positives and producing more accurate detection results. Our proposed framework can also be trained with weak supervision for those datasets that lack character-level annotations. Experiments on several datasets show that the proposed TextFuseNet achieves state-of-the-art performance. Specifically, we achieve an F-measure of 94.3% on ICDAR2013, 92.1% on ICDAR2015, 87.1% on Total-Text and 86.6% on CTW-1500, respectively.

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Scene Text DetectionText Detection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Scene Text Detection IC19-Art TextFuseNet (ResNeXt-101) H-Mean 78.6 #3 of 4 Archive leaderboard report
Scene Text Detection ICDAR 2013 TextFuseNet (ResNeXt-101) F-Measure 94.61% #1 of 16 Archive leaderboard report
Scene Text Detection ICDAR 2013 TextFuseNet (ResNeXt-101) Precision 97.27 #1 of 16 Archive leaderboard report
Scene Text Detection ICDAR 2013 TextFuseNet (ResNeXt-101) Recall 92.09 #1 of 16 Archive leaderboard report
Scene Text Detection ICDAR 2015 TextFuseNet (ResNeXt-101) F-Measure 92.23 #1 of 43 Archive leaderboard report
Scene Text Detection ICDAR 2015 TextFuseNet (ResNeXt-101) Precision 93.96 #1 of 43 Archive leaderboard report
Scene Text Detection ICDAR 2015 TextFuseNet (ResNeXt-101) Recall 90.56 #1 of 43 Archive leaderboard report
Scene Text Detection SCUT-CTW1500 TextFuseNet (ResNeXt-101) F-Measure 87.4 #4 of 17 Archive leaderboard report
Scene Text Detection SCUT-CTW1500 TextFuseNet (ResNeXt-101) Precision 89.7 #4 of 17 Archive leaderboard report
Scene Text Detection SCUT-CTW1500 TextFuseNet (ResNeXt-101) Recall 85.1 #4 of 17 Archive leaderboard report
Scene Text Detection Total-Text TextFuseNet (ResNeXt-101) F-Measure 87.5% #5 of 27 Archive leaderboard report
Scene Text Detection Total-Text TextFuseNet (ResNeXt-101) Precision 89.2 #5 of 27 Archive leaderboard report
Scene Text Detection Total-Text TextFuseNet (ResNeXt-101) Recall 85.8 #5 of 27 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections