Browse State-of-the-Art › Key Information Extraction
Key Information Extraction
40 papers with code · 6 benchmarks · 12 datasets archive 2025-07-28
Key Information Extraction (KIE) is aimed at extracting structured information (e.g. key-value pairs) from form-style documents (e.g. invoices), which makes an important step towards intelligent document understanding.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
6 leaderboard tables shown for this task, 6 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| CORD (9 rows) | RORE (GeoLayoutLM) | Modeling Layout Reading Order as Ordering Relations for... | code | — | Compare |
| SROIE (5 rows) | LayoutLMv2LARGE (Excluding OCR mismatch) | LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document... | code | — | Compare |
| Kleister NDA (3 rows) | LayoutLMv2LARGE | LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document... | code | — | Compare |
| EPHOIE (1 row) | LayoutLMv3 | LayoutLMv3: Pre-training for Document AI with Unified Text and... | code | — | Compare |
| ETD500 (1 row) | CRF-visual | Automatic Metadata Extraction Incorporating Visual Features from... | code | — | Compare |
| SIMARA (1 row) | DAN | SIMARA: a database for key-value information extraction from full pages | — | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
12 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 40 papers with code (74 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
31 Dec 2019 19 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 1 pointer-only (licence)In this paper, we propose the \textbf{LayoutLM} to jointly model interactions between text and layout information across scanned document images, which is beneficial for a great number of real-world document image…
-
29 Dec 2020 9 repositories listedPre-training of text and layout has proved effective in a variety of visually-rich document understanding tasks due to its effective model architecture and the advantage of large-scale unlabeled scanned/digital-born…
-
28 Feb 2022 5 repositories listed Syntology ran 2 of 3 samples · 1 unverifiedLiLT can be pre-trained on the structured documents of a single language and then directly fine-tuned on other languages with the corresponding off-the-shelf monolingual/multilingual pre-trained textual models.
-
18 Apr 2022 4 repositories listedIn this paper, we propose \textbf{LayoutLMv3} to pre-train multimodal Transformers for Document AI with unified text and image masking.
-
27 May 2024 2 repositories listedIn the domain of Document AI, parsing semi-structured image form is a crucial Key Information Extraction (KIE) task.
-
17 Oct 2023 2 repositories listedHowever, BIO-tagging scheme relies on the correct order of model inputs, which is not guaranteed in real-world NER on scanned VrDs where text are recognized and arranged by OCR systems.
-
12 Oct 2022 2 repositories listed Syntology ran 2 of 7 samples · 5 unverifiedRecent years have witnessed the rise and success of pre-training techniques in visually-rich document understanding.
-
10 Aug 2021 2 repositories listed Syntology ran 2 of 13 samples · 11 unverifiedOn the other hand, this paper tackles the problem by going back to the basic: effective combination of text and layout.
-
1 Jul 2021 2 repositories listedOur experiments show that CRF with visual features outperformed both a heuristic and a CRF model with only text-based features.
-
26 Mar 2021 2 repositories listedIn order to roundly evaluate our proposed method as well as boost the future research, we release a new dataset named WildReceipt, which is collected and annotated tailored for the evaluation of key information…
-
16 Apr 2020 2 repositories listed Syntology ran 2 of 6 samples · 4 unverifiedComputer vision with state-of-the-art deep learning models has achieved huge success in the field of Optical Character Recognition (OCR) including text detection and recognition tasks recently.
-
8 Jul 2025 1 repository listedThis technical report introduces PaddleOCR 3.
-
26 Jun 2025 1 repository listedDocument understanding and analysis have received a lot of attention due to their widespread application.
-
22 Feb 2025 1 repository listedIn this paper, we introduce OmniParser V2, a universal model that unifies VsTP typical tasks, including text spotting, key information extraction, table recognition, and layout analysis, into a unified framework.
-
2 Oct 2024 1 repository listedKey information extraction (KIE) from visually rich documents (VRD) has been a challenging task in document intelligence because of not only the complicated and diverse layouts of VRD that make the model hard to…
-
29 Sep 2024 1 repository listedHowever, we argue that this formulation does not adequately convey the complete reading order information in the layout, which may potentially lead to performance decline in downstream VrD tasks.
-
11 Sep 2024 1 repository listedThis paper presents a novel approach to information extraction (IE) from visually rich documents (VRD) by employing a directed weighted graph representation to capture relationships among various VRD components.
-
2 Jul 2024 1 repository listed Syntology ran 7 of 7 samples · 0 unverified · 7 pointer-only (licence)However, existing methods that integrate spatial layouts with text have limitations, such as producing overly long text sequences or failing to fully leverage the autoregressive traits of LLMs.
-
1 May 2024 1 repository listedIn recent years, the challenge of extracting information from business documents has emerged as a critical task, finding applications across numerous domains.
-
28 Mar 2024 1 repository listed Syntology ran 3 of 4 samples · 1 unverifiedRecently, visually-situated text parsing (VsTP) has experienced notable advancements, driven by the increasing demand for automated document understanding and the emergence of Generative Large Language Models (LLMs)…
-
7 Mar 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)We present TextMonkey, a large multimodal model (LMM) tailored for text-centric tasks.
-
2 Feb 2024 1 repository listedNamed Entity Recognition (NER) is a key information extraction task with a long-standing tradition.
-
7 Jan 2024 1 repository listedHowever, simply concatenating SER and RE serially can lead to severe error propagation, and it fails to handle cases like multi-line entities in real scenarios.
-
5 Jan 2024 1 repository listedThis paper introduces a novel system for information extraction from visually rich documents (VRD) using a weighted graph representation.
-
1 Jan 2024 1 repository listedRecently visually-situated text parsing (VsTP) has experienced notable advancements driven by the increasing demand for automated document understanding and the emergence of Generative Large Language Models (LLMs)…
-
25 Oct 2023 1 repository listed Syntology ran 4 of 4 samples · 0 unverified · 4 pointer-only (licence)We assess the model's performance across a range of OCR tasks, including scene text recognition, handwritten text recognition, handwritten mathematical expression recognition, table structure recognition, and…
-
24 Oct 2023 1 repository listedKey information extraction (KIE) from scanned documents has gained increasing attention because of its applications in various domains.
-
18 Sep 2023 1 repository listedIn this paper, we present AMuRD, a novel multilingual human-annotated dataset specifically designed for information extraction from receipts.
-
13 May 2023 1 repository listedIn this paper, we conducted a comprehensive evaluation of Large Multimodal Models, such as GPT4V and Gemini, in various text-related visual tasks including Text Recognition, Scene Text-Centric Visual Question Answering…
-
28 Apr 2023 1 repository listedAdvances in the Visually-rich Document Understanding (VrDU) field and particularly the Key-Information Extraction (KIE) task are marked with the emergence of efficient Transformer-based approaches such as the LayoutLM…
Syntology lines on 9 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections