Datasets › FUNSD

FUNSD (Form Understanding in Noisy Scanned Documents)

Introduced by Guillaume Jaume et al. in FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents archive 2025-07-28

Form Understanding in Noisy Scanned Documents (FUNSD) comprises 199 real, fully annotated, scanned forms. The documents are noisy and vary widely in appearance, making form understanding (FoUn) a challenging task. The proposed dataset can be used for various tasks, including text detection, optical character recognition, spatial layout analysis, and entity labeling/linking.

Source: FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents

Image source: https://guillaumejaume.github.io/FUNSD/

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Semantic entity labeling FUNSD LayoutMask (large) F1 93.20 LayoutMask: Enhance Text-Layout Interaction in... — 15 Compare
Relation Extraction FUNSD LayoutLMv3 large EM + BBO + RSF F1 90.81 A LayoutLMv3-Based Model for Enhanced Relation... — 9 Compare
Entity Linking FUNSD GeoLayoutLM F1 89.45 GeoLayoutLM: Geometric Pre-training for Visual... alibabaresearch/advancedliteratemachinery 7 Compare

Papers archive 2025-07-28

17 shown of 17 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 179. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Modeling Layout Reading Order as Ordering Relations for Visually-rich Document Understanding 1 3 29 Sep 2024 not harvested
A LayoutLMv3-Based Model for Enhanced Relation Extraction in Visually-Rich Documents 0 1 16 Apr 2024 not harvested
Reading Order Matters: Information Extraction from Visually-rich Documents by Token Path Prediction 2 3 17 Oct 2023 not harvested
DocTr: Document Transformer for Structured Information Extraction in Documents 0 2 16 Jul 2023 not harvested
LayoutMask: Enhance Text-Layout Interaction in Multi-modal Pre-training for Document Understanding 0 2 30 May 2023 not harvested
GeoLayoutLM: Geometric Pre-training for Visual Information Extraction 1 4 21 Apr 2023 not harvested
StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training 1 2 1 Mar 2023 not harvested
DGCN Based Solution for Entity Linking on Visual Rich Document 0 1 16 Nov 2022 not harvested
ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding 2 1 12 Oct 2022 ran 2 of 7 samples (5 unverified)
XDoc: Unified Pre-training for Cross-Format Document Understanding 1 1 6 Oct 2022 not harvested
Doc2Graph: a Task Agnostic Document Understanding Framework based on Graph Neural Networks 1 2 23 Aug 2022 not harvested
LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking 4 2 18 Apr 2022 not harvested
LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding 5 1 28 Feb 2022 ran 2 of 3 samples (1 unverified)
Entity Relation Extraction as Dependency Parsing in Visually Rich Documents 0 1 19 Oct 2021 not harvested
BROS: A Pre-trained Language Model Focusing on Text and Layout for Better Key Information Extraction from Documents 2 1 10 Aug 2021 ran 2 of 13 samples (11 unverified)
LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding 9 3 29 Dec 2020 not harvested
LayoutLM: Pre-training of Text and Layout for Document Image Understanding 19 1 31 Dec 2019 ran 2 of 3 samples (1 unverified; 1 pointer-only for licence)

Dataset loaders archive 2025-07-28

2 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • FUNSD

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections