Browse State-of-the-Art › Table Extraction
Table Extraction
15 papers with code · 0 benchmarks · 0 datasets archive 2025-07-28
Table extraction involves detecting and recognizing a table's logical structure and content from its unstructured presentation within a document
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
No dataset record in the archive lists this task.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
15 shown of 15 papers with code (32 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
6 Jan 2020 5 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedThis includes accurate detection of the tabular region within an image, and subsequently detecting and extracting information from the rows and columns of the detected table.
-
15 Nov 2022 3 repositories listedThe goals of this survey are to provide a profound comprehension of the major developments in the field of Table Detection, offer insight into the different methodologies, and provide a systematic taxonomy of the…
-
30 Sep 2021 2 repositories listed Syntology ran 5 of 13 samples · 8 unverified · 10 pointer-only (licence)We demonstrate that these improvements lead to a significant increase in training performance and a more reliable estimate of model performance at evaluation for table structure recognition.
-
17 Jun 2025 1 repository listedWe propose QUEST, a Quality-aware Semi-supervised Table extraction framework designed for business documents.
-
5 Dec 2024 1 repository listedTo demonstrate the effectiveness of our dataset in training models to extract information from table images, we create FinTabQA, a layout large language model trained on an extractive question-answering task.
-
8 Sep 2024 1 repository listedTo address these issues, we have introduced the PDF table extraction (PdfTable) toolkit.
-
29 Jun 2024 1 repository listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)Tabular reasoning involves interpreting natural language queries about tabular data, which presents a unique challenge of combining language understanding with structured data analysis.
-
22 Feb 2024 1 repository listedRecently, interpreting complex charts with logical reasoning has emerged as challenges due to the development of vision-language models.
-
23 May 2023 1 repository listedWe use this collection of annotated tables to evaluate the ability of open-source and API-based language models to extract information from tables covering diverse domains and data formats.
-
2 Feb 2023 1 repository listedWe define the task of Contextualized Table Extraction (CTE), which aims to extract and define the structure of tables considering the textual context of the document.
-
23 Aug 2022 1 repository listedTables are widely used in several types of documents since they can bring important information in a structured way.
-
3 Jul 2022 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)A crucial component in the curation of KB for a scientific domain (e.
-
ScanBank: A Benchmark Dataset for Figure Extraction from Scanned Electronic Theses and Dissertations23 Jun 2021 1 repository listedTo the best of our knowledge, ScanBank is the first manually annotated dataset for figure and table extraction for scanned ETDs.
-
25 May 2021 1 repository listedMoreover, to incorporate the extraction of semantic information, we develop a graph-based table interpretation method.
-
17 Mar 2020 1 repository listedTabular data is a crucial form of information expression, which can organize data in a standard structure for easy information retrieval and comparison.
Syntology lines on 4 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections