Datasets › RVL-CDIP

RVL-CDIP

Introduced by Adam W. Harley et al. in Evaluation of Deep Convolutional Nets for Document Image Classification and Retrieval1 Jan 2015 archive 2025-07-28

The RVL-CDIP dataset consists of scanned document images belonging to 16 classes such as letter, form, email, resume, memo, etc. The dataset has 320,000 training, 40,000 validation and 40,000 test images. The images are characterized by low quality, noise, and low resolution, typically 100 dpi.

Source: Towards a Multi-modal, Multi-task Learning based Pre-training Framework for Document Representation Learning Image Source: https://www.cs.cmu.edu/~aharley/rvl-cdip/

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Document Image Classification RVL-CDIP EAML Accuracy 97.70% EAML: Ensemble Self-Attention-based Mutual Learning... — 31 Compare
Document Layout Analysis RVL-CDIP VisualWordGrid FAR 28.7 VisualWordGrid: Information Extraction From Scanned... — 1 Compare

Papers archive 2025-07-28

25 shown of 25 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 107. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
DoPTA: Improving Document Layout Analysis using Patch-Text Alignment 0 1 17 Dec 2024 not harvested
GlobalDoc: A Cross-Modal Vision-Language Framework for Real-World Document Image Retrieval and Classification 0 1 11 Sep 2023 not harvested
EAML: Ensemble Self-Attention-based Mutual Learning Network for Document Image Classification 0 1 11 May 2023 not harvested
StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training 1 2 1 Mar 2023 not harvested
Multimodal Side-Tuning for Document Classification 1 2 16 Jan 2023 not harvested
VLCDoC: Vision-Language Contrastive Pre-Training Model for Cross-Modal Document Classification 0 1 24 May 2022 not harvested
LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking 4 2 18 Apr 2022 not harvested
DocXClassifier: High Performance Explainable Deep Network for Document Image Classification 1 1 17 Mar 2022 not harvested
DiT: Self-supervised Pre-training for Document Image Transformer 4 2 4 Mar 2022 ran 0 of 11 samples (11 unverified)
LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding 5 1 28 Feb 2022 ran 2 of 3 samples (1 unverified)
OCR-free Document Understanding Transformer 5 1 30 Nov 2021 ran 0 of 9 samples (9 unverified)
DocFormer: End-to-End Transformer for Document Understanding 1 2 22 Jun 2021 ran 5 of 10 samples (5 unverified)
BEiT: BERT Pre-Training of Image Transformers 14 1 15 Jun 2021 ran 6 of 11 samples (5 unverified)
LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding 6 1 18 Apr 2021 not harvested
Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer 1 2 18 Feb 2021 not harvested
LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding 9 2 29 Dec 2020 not harvested
Training data-efficient image transformers & distillation through attention 40 1 23 Dec 2020 ran 12 of 19 samples (7 unverified; 3 pointer-only for licence)
VisualWordGrid: Information Extraction From Scanned Documents Using A Multimodal Approach 0 1 5 Oct 2020 not harvested
Visual and Textual Deep Feature Fusion for Document Image Classification 0 1 16 Jun 2020 not harvested
Improving accuracy and speeding up Document Image Classification through parallel systems 1 1 16 Jun 2020 not harvested
LayoutLM: Pre-training of Text and Layout for Document Image Understanding 19 1 31 Dec 2019 ran 2 of 3 samples (1 unverified; 1 pointer-only for licence)
RoBERTa: A Robustly Optimized BERT Pretraining Approach 67 1 26 Jul 2019 ran 22 of 48 samples (26 unverified; 23 pointer-only for licence)
Document Image Classification with Intra-Domain Transfer Learning and Stacked Generalization of Deep Convolutional Neural Networks 4 1 29 Jan 2018 ran 1 of 1 samples (0 unverified)
Analysis of Convolutional Neural Networks for Document Image Classification 0 1 10 Aug 2017 not harvested
Cutting the Error by Half: Investigation of Very Deep CNN and Advanced Training Strategies for Document Image Classification 5 1 11 Apr 2017 not harvested

Dataset loaders archive 2025-07-28

5 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • RVL-CDIP
  • rvl_cdip

2 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections