Datasets › PubLayNet

PubLayNet

Introduced by Xu Zhong et al. in PubLayNet: largest dataset ever for document layout analysis archive 2025-07-28

PubLayNet is a dataset for document layout analysis by automatically matching the XML representations and the content of over 1 million PDF articles that are publicly available on PubMed Central. The size of the dataset is comparable to established computer vision datasets, containing over 360 thousand document images, where typical document layout elements are annotated.

Source: PubLayNet: largest dataset ever for document layout analysis

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Document Layout Analysis PubLayNet val VGT Overall 0.962 Vision Grid Transformer for Document Layout Analysis alibabaresearch/advancedliteratemachinery 15 Compare

Papers archive 2025-07-28

13 shown of 13 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 123. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
DoPTA: Improving Document Layout Analysis using Patch-Text Alignment 0 1 17 Dec 2024 not harvested
Vision Grid Transformer for Document Layout Analysis 1 2 29 Aug 2023 not harvested
A Graphical Approach to Document Layout Analysis 1 1 3 Aug 2023 not harvested
Bridging the Performance Gap between DETR and R-CNN for Graphical Object Detection in Document Images 0 1 23 Jun 2023 not harvested
Transformer-based Approach for Document Understanding 0 1 16 Oct 2022 not harvested
Unified Pretraining Framework for Document Understanding 0 1 22 Apr 2022 not harvested
LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking 4 1 18 Apr 2022 not harvested
DiT: Self-supervised Pre-training for Document Image Transformer 4 1 4 Mar 2022 ran 0 of 11 samples (11 unverified)
BEiT: BERT Pre-Training of Image Transformers 14 1 15 Jun 2021 ran 6 of 11 samples (5 unverified)
VSR: A Unified Framework for Document Layout Analysis combining Vision, Semantics and Relations 1 1 13 May 2021 not harvested
Training data-efficient image transformers & distillation through attention 40 1 23 Dec 2020 ran 12 of 19 samples (7 unverified; 3 pointer-only for licence)
CDeC-Net: Composite Deformable Cascade Network for Table Detection in Document Images 3 1 25 Aug 2020 ran 0 of 6 samples (6 unverified)
PubLayNet: largest dataset ever for document layout analysis 6 2 16 Aug 2019 ran 7 of 13 samples (6 unverified; 13 pointer-only for licence)

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • PubLayNet
  • PubLayNet val

2 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections