Browse State-of-the-Art › Document AI
Document AI
24 papers with code · 1 benchmark · 1 dataset archive 2025-07-28
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| EPHOIE (1 row) | LayoutLMv3 | LayoutLMv3: Pre-training for Document AI with Unified Text and... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
24 shown of 24 papers with code (40 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
31 Dec 2019 19 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 1 pointer-only (licence)In this paper, we propose the \textbf{LayoutLM} to jointly model interactions between text and layout information across scanned document images, which is beneficial for a great number of real-world document image…
-
5 Dec 2022 5 repositories listed Syntology ran 4 of 17 samples · 13 unverified · 3 pointer-only (licence)UDOP leverages the spatial correlation between textual content and document image to model image, text, and layout modalities with one uniform representation.
-
18 Apr 2022 4 repositories listedIn this paper, we propose \textbf{LayoutLMv3} to pre-train multimodal Transformers for Document AI with unified text and image masking.
-
4 Mar 2022 4 repositories listed Syntology ran 0 of 11 samples · 11 unverifiedWe leverage DiT as the backbone network in a variety of vision-based Document AI tasks, including document image classification, document layout analysis, table detection as well as text detection for OCR.
-
27 May 2024 2 repositories listedIn the domain of Document AI, parsing semi-structured image form is a crucial Key Information Extraction (KIE) task.
-
8 Apr 2024 2 repositories listed Syntology ran 1 of 2 samples · 1 unverifiedThe core of LayoutLLM is a layout instruction tuning strategy, which is specially designed to enhance the comprehension and utilization of document layouts.
-
18 Jul 2023 2 repositories listedWe address the extraction of mathematical statements and their proofs from scholarly PDF articles as a multimodal classification problem, utilizing text, font features, and bitmap image renderings of PDFs as distinct…
-
1 Jun 2025 1 repository listedAutomated parsing of scanned documents into richly structured, machine-readable formats remains a critical bottleneck in Document AI, as traditional multi-stage pipelines suffer from error propagation and limited…
-
20 Aug 2024 1 repository listedTo address the scarcity of high-quality publicly available document datasets and encourage further research on OOD detection for documents, we introduce FinanceDocs, a new document AI dataset.
-
8 Aug 2024 1 repository listedThe EU AI Act mandates that providers and deployers of high-risk AI systems establish a quality management system (QMS).
-
26 Jul 2024 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedOffice automation significantly enhances human productivity by automatically finishing routine tasks in the workflow.
-
17 Jun 2024 1 repository listed Syntology ran 6 of 8 samples · 2 unverified · 8 pointer-only (licence)Recent advancements in language and vision assistants have showcased impressive capabilities but suffer from a lack of transparency, limiting broader research and reproducibility.
-
7 May 2024 1 repository listedThis underscores the potential of DocRes across a broader spectrum of document image restoration tasks.
-
23 Oct 2023 1 repository listedThe use of visually-rich documents (VRDs) in various fields has created a demand for Document AI models that can read and comprehend documents like humans, which requires the overcoming of technical, linguistic, and…
-
19 Oct 2023 1 repository listedIn this report, we introduce DocXChain, a powerful open-source toolchain for document parsing, which is designed and developed to automatically convert the rich information embodied in unstructured documents, such as…
-
29 Aug 2023 1 repository listedDocument pre-trained models and grid-based models have proven to be very effective on various tasks in Document AI.
-
29 Aug 2023 1 repository listedIn this study, we aim to fill these gaps by conducting a comparative evaluation of state-of-the-art models in document layout analysis and investigating the potential of cross-lingual layout analysis by utilizing…
-
15 May 2023 1 repository listed Syntology ran 0 of 5 samples · 5 unverifiedWe call on the Document AI (DocAI) community to reevaluate current methodologies and embrace the challenge of creating more practically-oriented benchmarks.
-
7 May 2023 1 repository listedAs a prerequisite of chart data extraction, the accurate detection of chart basic elements is essential and mandatory.
-
21 Apr 2023 1 repository listedAdditionally, novel relation heads, which are pre-trained by the geometric pre-training tasks and fine-tuned for RE, are elaborately designed to enrich and enhance the feature representation.
-
9 Mar 2023 1 repository listedTo this end, we propose a simple but effective in-context learning framework called ICL-D3IE, which enables LLMs to perform DIE with different types of demonstration examples.
-
9 Nov 2022 1 repository listedAn initial document-specific model can be trained and its inference can be used as feedback for generating more automated annotations.
-
1 Jun 2022 1 repository listedThis paper was submitted for Financial Narrative Summarization (FNS) task in FNP-2022 workshop.
-
23 May 2022 1 repository listedThe processing of Visually-Rich Documents (VRDs) is highly important in information extraction tasks associated with Document Intelligence.
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections