Papers › Star-net: A spatial attention residue network for scene text recognition.

Star-net: A spatial attention residue network for scene text recognition.

20 Sep 2016The British Machine Vision Conference,2016 2016 9archive 2025-07-28

W. Liu, C. Chen, K.-Y. K. Wong, Z. Su, and J. Han.

In this paper, we present a novel SpaTial Attention Residue Network (STAR-Net) for recognising scene texts. Our STAR-Net is equipped with a spatial attention mechanism which employs a spatial transformer to remove the distortions of texts in natural images. This allows the subsequent feature extractor to focus on the rectified text region without being sidetracked by the distortions. Our STAR-Net also exploits residue convolutional blocks to build a very deep feature extractor, which is essential to the successful extraction of discriminative text features for this fine grained recognition task. Combining the spatial attention mechanism with the residue convolutional blocks, our STAR-Net is the deepest end-to-end trainable neural network for scene text recognition. Experiments have been conducted on five public benchmark datasets. Experimental results show that our STAR-Net can achieve a performance comparable to state-of-the-art methods for scene texts with little distortions, and outperform these methods for scene texts with considerable distortions.

PaperPDFCode

Code

PaddlePaddle/PaddleOCR paddleApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Optical Character Recognition (OCR)Scene Text Recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Scene Text Recognition ICDAR 2003 STAR-Net Accuracy 89.9 #11 of 12 Archive leaderboard report
Scene Text Recognition ICDAR2013 STAR-Net Accuracy 89.1 #35 of 38 Archive leaderboard report
Scene Text Recognition SVT STAR-Net Accuracy 83.6 #34 of 37 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Spatial Transformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections