Home › Datasets › task › Document Layout Analysis

Document Layout Analysis datasets

archive 2025-07-28

13 datasets carry the task tag "Document Layout Analysis" (the task itself: Document Layout Analysis), ordered by the archive's paper count. Page 1 of 1: 13 shown of 13. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Document Layout Analysis datasets 1–13 of 13

PubLayNet is a dataset for document layout analysis by automatically matching the XML representations and the content of over 1 million PDF articles that are publicly available on PubMed Central.
123 papers · 1 benchmark
The RVL-CDIP dataset consists of scanned document images belonging to 16 classes such as letter, form, email, resume, memo, etc.
107 papers · 2 benchmarks
A benchmark dataset that contains 500K document pages with fine-grained token-level annotations for document layout analysis.
34 papers · 0 benchmarks
The database consists of 150 annotated pages of three different medieval manuscripts with challenging layouts.
15 papers · 2 benchmarks
The DSSE-200 is a complex document layout dataset including various dataset styles.
8 papers · 0 benchmarks
HJDataset is a large dataset of Historical Japanese Documents with Complex Layouts.
5 papers · 0 benchmarks
We present the VIS30K dataset, a collection of 29,689 images that represents 30 years of figures and tables from each track of the IEEE Visualization conference series (Vis, SciVis, InfoVis, VAST).
5 papers · 0 benchmarks
The D4LA dataset is a diverse benchmark for document layout analysis (DLA) derived from the RVL-CDIP dataset.
3 papers · 1 benchmark
U-DIADS-Bib is a proprietary dataset developed through the collaboration of computer scientists and humanities at the University of Udine.
2 papers · 1 benchmark
- Revision: v1.0.0-full-20210527a - DOI: 10.5281/zenodo.4817662 - Authors: J.
1 paper · 0 benchmarks
We compiled a new dataset (the PERO layout dataset) that contains 683 images from various sources and historical periods with complete manual text block, text line polygon and baseline annotations.
1 paper · 0 benchmarks
The UrduDoc Dataset is a benchmark dataset for Urdu text line detection in scanned documents.
1 paper · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.