Browse State-of-the-Art › document understanding › Papers, page 2
document understanding
Papers archive 2025-07-28
archive papers tagged: 309 · with a code link: 140 · where Syntology ran a sample: 39 (30 with a run with no instrument failure, 9 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (39 of 309 tagged: 30 with a run with no instrument failure, 9 where every run was a failure of Syntology's instrument)
Page 2 of 4: papers 101 to 200 of 309, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
3 May 2023 1 repository listed
-
28 Apr 2023 1 repository listed
-
28 Apr 2023 1 repository listed
-
23 Mar 2023 1 repository listed
-
1 Jan 2023 1 repository listed
-
19 Dec 2022 1 repository listed
-
6 Dec 2022 1 repository listed
-
7 Nov 2022 1 repository listed
-
8 Oct 2022 1 repository listed Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
6 Oct 2022 1 repository listed
-
1 Oct 2022 1 repository listed
-
23 Aug 2022 1 repository listed
-
28 Jul 2022 1 repository listed
-
14 Jul 2022 1 repository listed
-
1 Jun 2022 1 repository listed
-
1 May 2022 1 repository listed
-
1 May 2022 1 repository listed
-
25 Mar 2022 1 repository listed
-
14 Mar 2022 1 repository listed
-
16 Jan 2022 1 repository listed
-
25 Dec 2021 1 repository listed
-
15 Dec 2021 1 repository listed
-
2 Sep 2021 1 repository listed
-
22 Jun 2021 1 repository listed Syntology 5 ran (of which 5 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified; every one of the 5 samples that ran constructed an object rather than computing a result (of 10 harvested samples)
-
23 May 2021 1 repository listed
-
18 Feb 2021 1 repository listed
-
24 Jan 2021 1 repository listed
-
18 Dec 2020 1 repository listed
-
14 Dec 2020 1 repository listed
-
9 Dec 2020 1 repository listed
-
7 Dec 2020 1 repository listed
-
27 Nov 2020 1 repository listed
-
3 Nov 2020 1 repository listed
-
22 Oct 2020 1 repository listed Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
12 Oct 2020 1 repository listed
-
27 May 2020 1 repository listed
-
7 Nov 2019 1 repository listed
-
25 Oct 2019 1 repository listed
-
16 Aug 2019 1 repository listed
-
21 Feb 2018 1 repository listed
-
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends14 Jul 2025 0 repositories listed
-
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models25 Jun 2025 0 repositories listed
-
WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts18 Jun 2025 0 repositories listed
-
A Survey on Vietnamese Document Analysis and Recognition: Challenges and Future Directions5 Jun 2025 0 repositories listed
-
DiCoRe: Enhancing Zero-shot Event Detection via Divergent-Convergent LLM Reasoning5 Jun 2025 0 repositories listed
-
MT³: Scaling MLLM-based Text Image Machine Translation via Multi-Task Reinforcement Learning26 May 2025 0 repositories listed
-
Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning26 May 2025 0 repositories listed
-
Doc-CoB: Enhancing Multi-Modal Document Understanding with Visual Chain-of-Boxes Reasoning24 May 2025 0 repositories listed
-
The Hidden Structure -- Improving Legal Document Understanding Through Explicit Text Formatting19 May 2025 0 repositories listed
-
WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild?16 May 2025 0 repositories listed
-
Document Image Rectification Bases on Self-Adaptive Multitask Fusion9 May 2025 0 repositories listed
-
Automated Parsing of Engineering Drawings for Structured Information Extraction Using a Fine-tuned Document Understanding Transformer2 May 2025 0 repositories listed
-
NoTeS-Bank: Benchmarking Neural Transcription and Search for Scientific Notes Understanding12 Apr 2025 0 repositories listed
-
QID: Efficient Query-Informed ViTs in Data-Scarce Regimes for OCR-free Visual Document Understanding3 Apr 2025 0 repositories listed
-
How does Watermarking Affect Visual Language Models in Document Understanding?1 Apr 2025 0 repositories listed
-
Improving Applicability of Deep Learning based Token Classification models during Training28 Mar 2025 0 repositories listed
-
M-DocSum: Do LVLMs Genuinely Comprehend Interleaved Image-Text in Document Summarization?27 Mar 2025 0 repositories listed
-
A Simple yet Effective Layout Token in Large Language Models for Document Understanding24 Mar 2025 0 repositories listed
-
18 Mar 2025 0 repositories listed Syntology 3 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 3 samples that ran constructed an object rather than computing a result (of 3 harvested samples) · 3 pointer-only (licence)
-
4 Mar 2025 0 repositories listed Syntology 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 6 harvested samples) · 6 pointer-only (licence)
-
Shakti-VLMs: Scalable Vision-Language Models for Enterprise AI24 Feb 2025 0 repositories listed
-
KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding20 Feb 2025 0 repositories listed
-
Assessing Generative AI value in a public sector context: evidence from a field experiment13 Feb 2025 0 repositories listed
-
BoundingDocs: a Unified Dataset for Document Question Answering with Spatial Annotations6 Jan 2025 0 repositories listed
-
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends4 Jan 2025 0 repositories listed
-
Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark1 Jan 2025 0 repositories listed
-
Zero-Shot Prompting and Few-Shot Fine-Tuning: Revisiting Document Image Classification Using Large Language Models18 Dec 2024 0 repositories listed
-
Memory-Augmented Agent Training for Business Document Understanding17 Dec 2024 0 repositories listed
-
DocVLM: Make Your VLM an Efficient Reader11 Dec 2024 0 repositories listed
-
BigDocs: An Open and Permissively-Licensed Dataset for Training Multimodal Models on Document and Code Tasks5 Dec 2024 0 repositories listed
-
MATATA: Weakly Supervised End-to-End MAthematical Tool-Augmented Reasoning for Tabular Applications28 Nov 2024 0 repositories listed
-
DOGE: Towards Versatile Visual Document Grounding and Referring26 Nov 2024 0 repositories listed
-
StructFormer: Document Structure-based Masked Attention and its Impact on Language Model Pre-Training25 Nov 2024 0 repositories listed
-
Information Extraction from Heterogeneous Documents without Ground Truth Labels using Synthetic Label Generation and Knowledge Distillation22 Nov 2024 0 repositories listed
-
Is Cognition consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding12 Nov 2024 0 repositories listed
-
M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework9 Nov 2024 0 repositories listed
-
Hierarchical Visual Feature Aggregation for OCR-Free Document Understanding8 Nov 2024 0 repositories listed
-
M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding7 Nov 2024 0 repositories listed
-
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection5 Nov 2024 0 repositories listed
-
LoRA-Contextualizing Adaptation of Large Multimodal Models for Long Document Understanding2 Nov 2024 0 repositories listed
-
MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding25 Oct 2024 0 repositories listed
-
"What is the value of {templates}?" Rethinking Document Information Extraction Datasets for LLMs20 Oct 2024 0 repositories listed
-
Harnessing Webpage UIs for Text-Rich Visual Understanding17 Oct 2024 0 repositories listed
-
ReLayout: Towards Real-World Document Understanding via Layout-enhanced Pre-training14 Oct 2024 0 repositories listed
-
DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models4 Oct 2024 0 repositories listed
-
DAViD: Domain Adaptive Visually-Rich Document Understanding with Synthetic Insights2 Oct 2024 0 repositories listed
-
Leveraging Long-Context Large Language Models for Multi-Document Understanding and Summarization in Enterprise Applications27 Sep 2024 0 repositories listed
-
DocMamba: Efficient Document Pre-training with State Space Model18 Sep 2024 0 repositories listed
-
Leveraging Distillation Techniques for Document Understanding: A Case Study with FLAN-T517 Sep 2024 0 repositories listed
-
ViRED: Prediction of Visual Relations in Engineering Drawings2 Sep 2024 0 repositories listed
-
The MERIT Dataset: Modelling and Efficiently Rendering Interpretable Transcripts31 Aug 2024 0 repositories listed
-
SynthDoc: Bilingual Documents Synthesis for Visual Document Understanding27 Aug 2024 0 repositories listed
-
Building and better understanding vision-language models: insights and future directions22 Aug 2024 0 repositories listed
-
Arctic-TILT. Business Document Understanding at Sub-Billion Scale8 Aug 2024 0 repositories listed
-
Deep Learning based Visually Rich Document Content Understanding: A Survey2 Aug 2024 0 repositories listed
-
Deep Learning based Key Information Extraction from Business Documents: Systematic Literature Review23 Jul 2024 0 repositories listed
-
16 Jul 2024 0 repositories listed
-
DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming27 Jun 2024 0 repositories listed
-
DrVideo: Document Retrieval Based Long Video Understanding18 Jun 2024 0 repositories listed
-
Enhancing Question Answering on Charts Through Effective Pre-training Tasks14 Jun 2024 0 repositories listed
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.