Browse State-of-the-Art › Scene Text Recognition
Scene Text Recognition
146 papers with code · 15 benchmarks · 29 datasets archive 2025-07-28
See Scene Text Detection for leaderboards in this task.
Description from the archive archive 2025-07-28; Papers-with-Code links inside it are rewritten to this site.
Benchmarks archive 2025-07-28
15 leaderboard tables shown for this task, 15 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 15 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
29 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 146 papers with code (269 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
21 Jul 2015 85 repositories listed Syntology ran 19 of 81 samples · 62 unverified · 17 pointer-only (licence)In this paper, we investigate the problem of scene text recognition, which is among the most important and challenging tasks in image-based sequence recognition.
-
3 Apr 2019 13 repositories listed Syntology ran 0 of 19 samples · 19 unverifiedMany new proposals for scene text recognition (STR) models have been introduced in recent years.
-
21 Sep 2021 8 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedText recognition is a long-standing research problem for document digitalization.
-
2 Nov 2018 8 repositories listed Syntology ran 0 of 8 samples · 8 unverifiedRecognizing irregular text in natural scene images is challenging due to the large variance in text appearance, such as curvature, orientation and distortion.
-
7 Oct 2019 7 repositories listed Syntology ran 4 of 10 samples · 6 unverifiedAttention-based scene text recognizers have gained huge success, which leverages a more compact intermediate representation to learn 1d- or 2d- attention by a RNN-based encoder-decoder architecture.
-
10 Jan 2019 7 repositories listedIt decreases the difficulty of recognition and enables the attention-based sequence recognition network to more easily read irregular text.
-
5 Jan 2018 7 repositories listedIncidental scene text spotting is considered one of the most difficult and valuable challenges in the document analysis community.
-
22 Aug 2021 5 repositories listedSuch operation guides the vision model to use not only the visual texture of characters, but also the linguistic information in visual context for recognition when the visual cues are confused (e.
-
15 Jul 2020 5 repositories listed Syntology ran 1 of 4 samples · 3 unverifiedTheoretically, our proposed method, dubbed \emph{RobustScanner}, decodes individual characters with dynamic ratio between context and positional clues, and utilizes more positional ones when the decoding sequences with…
-
21 Dec 2019 5 repositories listedTo remedy this issue, we propose a decoupled attention network (DAN), which decouples the alignment operation from using historical decoding results.
-
12 Mar 2016 5 repositories listedWe show that the model is able to recognize several types of irregular text, including perspective text and curved text.
-
30 Apr 2022 4 repositories listedDominant scene text recognition models commonly contain two building blocks, a visual model for feature extraction and a sequence model for text transcription.
-
10 May 2021 4 repositories listedIn this paper, we propose a primitive representation learning method that aims to exploit intrinsic representations of scene text images.
-
11 Mar 2021 4 repositories listed Syntology ran 18 of 28 samples · 10 unverified · 18 pointer-only (licence)Additionally, based on the ensemble of iterative predictions, we propose a self-training method which can learn from unlabeled images effectively.
-
25 Jun 2018 4 repositories listedSCENE text recognition has attracted great interest from the academia and the industry in recent years owing to its importance in a wide range of applications.
-
4 Jun 2018 4 repositories listedConsidering scene image has large variation in text and background, we further design a modality-transform block to effectively transform 2D input images to 1D sequences, combined with the encoder to extract more…
-
8 Sep 2022 3 repositories listedIn this work, we first draw inspiration from the recent progress in Vision Transformer (ViT) to construct a conceptually simple yet powerful vision STR model, which is built upon ViT and outperforms previous…
-
Multi-modal Text Recognition Networks: Interactive Enhancements between Visual and Semantic Features30 Nov 2021 3 repositories listedFurthermore, MATRN stimulates combining semantic features into visual features by hiding visual clues related to the character in the training phase.
-
22 Nov 2021 3 repositories listedIn this paper, we propose a novel module called Multi-Domain Character Distance Perception (MDCDP) to establish a visually and semantically related position embedding.
-
18 May 2021 3 repositories listedOn a comparable strong baseline method such as TRBA with accuracy of 84.
-
27 May 2020 3 repositories listedArbitrary text appearance poses a great challenge in scene text recognition tasks.
-
22 May 2020 3 repositories listedScene text recognition is a hot research topic in computer vision.
-
27 Mar 2020 3 repositories listedScene text image contains two levels of contents: visual texture and semantic information.
-
24 Mar 2020 3 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedSynthetic data has been a critical tool for training scene text detection and recognition models.
-
29 Oct 2018 3 repositories listedWe propose a post-processing approach to improve scene text recognition accuracy by using occurrence probabilities of words (unigram language model), and the semantic correlation between scene and text.
-
9 Jan 2018 3 repositories listedIn this paper, we present an end-to-end trainable fast scene text detector, named TextBoxes++, which detects arbitrary-oriented scene text with both high accuracy and efficiency in a single network forward pass.
-
27 Jul 2017 3 repositories listedIn contrast to most existing works that consist of multiple deep neural networks and several pre-processing steps we propose to use a single deep neural network that learns to detect and recognize text from natural…
-
21 Feb 2024 2 repositories listedBy enhancing the alignment between the canonical mask feature and the text feature, the module ensures more effective fusion, ultimately leading to improved recognition performance.
-
26 Jan 2024 2 repositories listed Syntology ran 4 of 11 samples · 7 unverifiedLeveraging the global structure of the text as a prior, the proposed GSDM develops an efficient diffusion model to recover clean texts.
-
25 Jul 2023 2 repositories listedSpecifically, MGP-STR achieves an average recognition accuracy of 94% on standard benchmarks for scene text recognition.
Syntology lines on 9 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections