Home › Datasets › task › Optical Character Recognition (OCR)

Optical Character Recognition (OCR) datasets

archive 2025-07-28

55 datasets carry the task tag "Optical Character Recognition (OCR)" (the task itself: Optical Character Recognition (OCR)), ordered by the archive's paper count. Page 1 of 2: 48 shown of 55. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Optical Character Recognition (OCR) datasets 1–48 of 55

IAM (IAM Handwriting)
The IAM database contains 13,353 images of handwritten lines of text created by 657 writers.
198 papers · 1 benchmark
FUNSD (Form Understanding in Noisy Scanned Documents)
Form Understanding in Noisy Scanned Documents (FUNSD) comprises 199 real, fully annotated, scanned forms.
179 papers · 3 benchmarks
Contains 145k captions for 28k images.
98 papers · 1 benchmark
ST-VQA (Scene Text Visual Question Answering)
ST-VQA aims to highlight the importance of exploiting high-level semantic information present in images as textual cues in the VQA process.
90 papers · 0 benchmarks
The ICDAR2003 dataset is a dataset for scene text recognition.
53 papers · 1 benchmark
TextOCR is a dataset to benchmark text recognition on arbitrary shaped scene-text.
37 papers · 0 benchmarks
SciTSR is a large-scale table structure recognition dataset, which contains 15,000 tables in PDF format and their corresponding structure labels obtained from LaTeX source files.
36 papers · 0 benchmarks
A benchmark dataset that contains 500K document pages with fine-grained token-level annotations for document layout analysis.
34 papers · 0 benchmarks
Description: 105,941 Images Natural Scenes OCR Data of 12 Languages.
24 papers · 0 benchmarks
This dataset includes 4,500 fully annotated images (over 30,000 license plate characters) from 150 vehicles in real-world scenarios where both the vehicle and the camera (inside another vehicle) are moving.
15 papers · 1 benchmark
A prebuilt dataset for OpenAI's task for image-2-latex system.
12 papers · 1 benchmark
MRR-Benchmark (Multi-Modal Reading Benchmark)
Multi-Modal Reading (MMR) Benchmark includes 550 annotated question-answer pairs across 11 distinct tasks involving texts, fonts, visual elements, bounding boxes, spatial relations, and grounding, with carefully designed evaluation metrics.
11 papers · 1 benchmark
This dataset, called RodoSol-ALPR dataset, contains 20,000 images captured by static cameras located at pay tolls owned by the Rodovia do Sol (RodoSol) concessionaire, which operates 67.5 kilometers of a highway (ES-060) in the Brazilian…
8 papers · 0 benchmarks
This dataset aims at evaluating the License Plate Character Segmentation (LPCS) problem.
8 papers · 1 benchmark
The Kannada-MNIST dataset is a drop-in substitute for the standard MNIST dataset for the Kannada language.
7 papers · 0 benchmarks
IIIT-AR-13K is created by manually annotating the bounding boxes of graphical or page objects in publicly available annual reports.
6 papers · 0 benchmarks
Contains video clips shot with modern high-resolution mobile cameras, with strong projective distortions and with low lighting conditions.
6 papers · 0 benchmarks
MLe2 is a dataset for the evaluation of scene text end-to-end reading systems and all intermediate stages such as text detection, script identification and text recognition.
6 papers · 0 benchmarks
Chinese Text in the Wild is a dataset of Chinese text with about 1 million Chinese characters from 3850 unique ones annotated by experts in over 30000 street view images.
5 papers · 0 benchmarks
We collect, organize and open-source the large-scale multimodal instruction dataset, Infinity-MM, consisting of tens of millions of samples.
5 papers · 0 benchmarks
Twitter100k is a large-scale dataset for weakly supervised cross-media retrieval.
5 papers · 0 benchmarks
The ChineseLP dataset contains 411 vehicle images (mostly of passenger cars) with Chinese license plates (LPs).
4 papers · 1 benchmark
DDI-100 (Distorted Document Images)
The DDI-100 dataset is a synthetic dataset for text detection and recognition based on 7000 real unique document pages and consists of more than 100000 augmented images.
4 papers · 0 benchmarks
This dataset contains Bangla handwritten numerals, basic characters and compound characters.
3 papers · 2 benchmarks
Arabic handwriting dataset.
3 papers · 1 benchmark
dataset link : https://www.kaggle.com/datasets/osamahosamabdellatif/high-quality-invoice-images-for-ocr Overview High-Quality Invoice Images for OCR is a curated dataset containing professionally scanned and digitally captured invoice…
3 papers · 0 benchmarks
Imgur5k is a large-scale handwritten in-the-wild dataset, containing challenging real world handwritten samples from nearly 5K writers.
3 papers · 0 benchmarks
The largest dataset of extracted visual content from historic newspapers ever produced.
3 papers · 0 benchmarks
MCSCSet is a large-scale specialist-annotated dataset, designed for the task of Medical-domain Chinese Spelling Correction that contains about 200k samples.
2 papers · 0 benchmarks
MSDA (Multi-source domain adaptation dataset for text recognition)
5 domains: synthetic domain, document domain, street view domain, handwritten domain, and car license domain over five million images
2 papers · 2 benchmarks
A large scale OCSR dataset, proposed in paper “MolParser: End-to-end Visual Recognition of Molecule Structures in the Wild“ MolParser-7M contains nearly 8 million paired image-SMILES data.
2 papers · 0 benchmarks
This dataset contains 2,000 images taken from inside a warehouse of the Energy Company of Paraná (Copel), which directly serves more than 4 million consuming units in the Brazilian state of Paraná.
2 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Data collection: Finding a suitable source of data is considered a first step toward building a database.
1 paper · 1 benchmark
BLN600 (BLN600: A Parallel Corpus of Machine/Human Transcribed Nineteenth Century Newspaper Texts)
A publicly available corpus of nineteenth-century newspaper text focused on crime in London, derived from the Gale British Library Newspapers corpus parts 1 and 2.
1 paper · 0 benchmarks
This dataset contains 12,500 meter images acquired in the field by the employees of the Energy Company of Paraná (Copel), which directly serves more than 4 million consuming units, across 395 cities and 1,113 locations (i.e., districts,…
1 paper · 1 benchmark
Doc3DShade extends Doc3D with realistic lighting and shading.
1 paper · 0 benchmarks
Optical images of printed circuit boards as well as detailed annotations of any text, logos, and surface-mount devices (SMDs).
1 paper · 0 benchmarks
Introduced by Singh, Sumeet S..
1 paper · 1 benchmark
IllusionChartest Dataset Characteristics IllusionChartest is a generated dataset containing 3,300 samples of images that feature sequences of 3 to 5 random characters.
1 paper · 0 benchmarks
It is composed of around 770k of color 256x256 RGB images extracted from the European Union Intellectual Property Office (EUIPO) open registry.
1 paper · 1 benchmark
MatriVasha: (MatriVasha: Compound Character atasetD)
MatriVasha the largest dataset of handwritten Bangla compound characters for research on handwritten Bangla compound character recognition.
1 paper · 0 benchmarks
NCSE v2.0 (NCSE v2.0: A Dataset of OCR-Processed 19th Century English Newspapers)
The NCSE v2.0 is a digitized collection of six 19th-century English periodicals The ground truth contains 358 cropped images of text blocks from 31 pages of 19th century newspaper data
1 paper · 0 benchmarks
PsOCR (Pashto OCR Dataset)
PsOCR is a large-scale synthetic dataset for Optical Character Recognition in low-resource Pashto language.
1 paper · 0 benchmarks
SUT (SUT: a new multi-purpose synthetic dataset for Farsi document image analysis)
This paper introduces a new large-scale dataset for Farsi document images, named SUT, which aims to tackle the challenges associated with obtaining diverse and substantial ground-truth data for supervised models in document image analysis…
1 paper · 2 benchmarks
The UTRSet-Real dataset is a comprehensive, manually annotated dataset specifically curated for Printed Urdu OCR research.
1 paper · 0 benchmarks
The UTRSet-Synth dataset is introduced as a complementary training resource to the UTRSet-Real Dataset, specifically designed to enhance the effectiveness of Urdu OCR models.
1 paper · 0 benchmarks
Dataset Introduction This dataset leverages VideoDB's Public Collection to offer a diverse range of videos featuring text-containing scenes.
1 paper · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.