Browse State-of-the-Art › Optical Character Recognition (OCR)
Optical Character Recognition (OCR)
462 papers with code · 6 benchmarks · 55 datasets archive 2025-07-28
Optical Character Recognition or Optical Character Reader (OCR) is the electronic or mechanical conversion of images of typed, handwritten or printed text into machine-encoded text, whether from a scanned document, a photo of a document, a scene-photo (for example the text on signs and billboards in a landscape photo, license plates in cars...) or from subtitle text superimposed on an image (for example: from a television broadcast)
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
6 leaderboard tables shown for this task, 6 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Benchmarking Chinese Text Recognition: Datasets, Baselines, and an Empirical Study (7 rows) | DTrOCR | DTrOCR: Decoder-only Transformer for Optical Character Recognition | code | — | Compare |
| VideoDB's OCR Benchmark Public Collection (5 rows) | GPT-4o | Benchmarking Vision-Language Models on Optical Character... | code | Syntology ran 0 of 5 samples · 5 unverified | Compare |
| FSNS - Test (3 rows) | AttentionOCR_Inception-resnet-v2_Location | Attention-based Extraction of Structured Information from Street... | code | — | Compare |
| SUT (2 rows) | Tesseract | SUT: a new multi-purpose synthetic dataset for Farsi document... | code | — | Compare |
| I2L-140K (2 rows) | I2L-NOPOOL | Teaching Machines to Code: Neural Markup Generation with Visual Attention | code | — | Compare |
| im2latex-100k (1 row) | I2L-STRIPS | Teaching Machines to Code: Neural Markup Generation with Visual Attention | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
55 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 55 until expanded.
Subtasks archive 2025-07-28
10 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 462 papers with code (1,243 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
21 Jul 2015 85 repositories listed Syntology ran 19 of 81 samples · 62 unverified · 17 pointer-only (licence)In this paper, we investigate the problem of scene text recognition, which is among the most important and challenging tasks in image-based sequence recognition.
-
11 Apr 2017 31 repositories listed Syntology ran 5 of 6 samples · 1 unverified · 1 pointer-only (licence)Previous approaches for scene text detection have already achieved promising performances across various benchmarks.
-
28 Mar 2019 19 repositories listedDue to the fact that there are large geometrical margins among the minimal scale kernels, our method is effective to split the close text instances, making it easier to use segmentation-based methods to detect…
-
20 Nov 2019 15 repositories listed Syntology ran 3 of 25 samples · 22 unverifiedRecently, segmentation-based methods are quite popular in scene text detection, as the segmentation results can more accurately describe scene text of various shapes such as curve text.
-
16 Sep 2016 14 repositories listed Syntology ran 1 of 13 samples · 12 unverifiedWe present a neural encoder-decoder model to convert images into presentational markup based on a scalable coarse-to-fine attention mechanism.
-
21 Sep 2020 10 repositories listed Syntology ran 0 of 12 samples · 12 unverifiedMeanwhile, several pre-trained models for the Chinese and English recognition are released, including a text detector (97K images are used), a direction classifier (600K images are used) as well as a text recognizer (17.
-
21 Sep 2020 10 repositories listed Syntology ran 0 of 12 samples · 12 unverifiedMeanwhile, several pre-trained models for the Chinese and English recognition are released, including a text detector (97K images are used), a direction classifier (600K images are used) as well as a text recognizer (17.
-
21 Sep 2021 8 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedText recognition is a long-standing research problem for document digitalization.
-
21 Sep 2021 8 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedText recognition is a long-standing research problem for document digitalization.
-
2 Nov 2018 8 repositories listed Syntology ran 0 of 8 samples · 8 unverifiedRecognizing irregular text in natural scene images is challenging due to the large variance in text appearance, such as curvature, orientation and distortion.
-
10 Jan 2019 7 repositories listedIt decreases the difficulty of recognition and enables the attention-based sequence recognition network to more easily read irregular text.
-
25 Nov 2019 6 repositories listed Syntology ran 2 of 6 samples · 4 unverifiedIn addition, we propose a new Tree-Edit-Distance-based Similarity (TEDS) metric for table recognition, which more appropriately captures multi-hop cell misalignment and OCR errors than the pre-established metric.
-
30 Nov 2021 5 repositories listed Syntology ran 0 of 9 samples · 9 unverifiedCurrent Visual Document Understanding (VDU) methods outsource the task of reading text to off-the-shelf Optical Character Recognition (OCR) engines and focus on the understanding task with the OCR outputs.
-
30 Nov 2021 5 repositories listed Syntology ran 0 of 9 samples · 9 unverifiedCurrent Visual Document Understanding (VDU) methods outsource the task of reading text to off-the-shelf Optical Character Recognition (OCR) engines and focus on the understanding task with the OCR outputs.
-
28 Jan 2021 5 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedInspired by the recent advance in unsupervised contrastive representation learning, we propose a pixel-wise contrastive framework for semantic segmentation in the fully supervised setting.
-
28 Feb 2018 5 repositories listed[python3.
-
12 Mar 2016 5 repositories listedWe show that the model is able to recognize several types of irregular text, including perspective text and curved text.
-
7 Oct 2022 4 repositories listed Syntology ran 0 of 5 samples · 5 unverifiedVisually-situated language is ubiquitous -- sources range from textbooks with diagrams to web pages with images and tables, to mobile apps with buttons and forms.
-
4 Mar 2022 4 repositories listed Syntology ran 0 of 11 samples · 11 unverifiedWe leverage DiT as the backbone network in a variety of vision-based Document AI tasks, including document image classification, document layout analysis, table detection as well as text detection for OCR.
-
17 Oct 2020 4 repositories listedDocuments often exhibit various forms of degradation, which make it hard to be read and substantially deteriorate the performance of an OCR system.
-
25 Jun 2018 4 repositories listedSCENE text recognition has attracted great interest from the academia and the industry in recent years owing to its importance in a wide range of applications.
-
25 Jun 2018 4 repositories listedSCENE text recognition has attracted great interest from the academia and the industry in recent years owing to its importance in a wide range of applications.
-
4 Jun 2018 4 repositories listedConsidering scene image has large variation in text and background, we further design a modality-transform block to effectively transform 2D input images to 1D sequences, combined with the encoder to extract more…
-
13 Feb 2017 4 repositories listedWe introduce the French Street Name Signs (FSNS) Dataset consisting of more than a million images of street name signs cropped from Google Street View images of France.
-
26 Jan 2016 4 repositories listedThe goal of COCO-Text is to advance state-of-the-art in text detection and recognition in natural images.
-
26 Jan 2016 4 repositories listedThe goal of COCO-Text is to advance state-of-the-art in text detection and recognition in natural images.
-
29 Sep 2024 3 repositories listedThis paper shows that fine-tuning a language model on synthetic data using an LM and using a character level Markov corruption process can significantly improve the ability to correct OCR errors.
-
10 Aug 2024 3 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedWith support of over 300+ LLMs and 50+ MLLMs, SWIFT stands as the open-source framework that provide the most comprehensive support for fine-tuning large models.
-
10 Aug 2024 3 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedWith support of over 300+ LLMs and 50+ MLLMs, SWIFT stands as the open-source framework that provide the most comprehensive support for fine-tuning large models.
-
1 Dec 2023 3 repositories listedDocument image dewarping is a crucial task in computer vision with numerous practical applications.
Syntology lines on 17 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections