Browse State-of-the-Art › Table annotation
Table annotation
23 papers with code · 0 benchmarks · 10 datasets archive 2025-07-28
Table annotation is the task of annotating a table with terms/concepts from knowledge graph or database schema. Table annotation is typically broken down into the following five subtasks:
- Cell Entity Annotation (CEA)
- Column Type Annotation (CTA)
- Column Property Annotation (CPA)
- Table Type Detection
- Row Annotation
The SemTab challenge is closely related to the Table Annotation problem. It is a yearly challenge which focuses on the first three tasks of table annotation and its purpose is to benchmark different table annotation systems.
Description from the archive archive 2025-07-28; Papers-with-Code links inside it are rewritten to this site.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
10 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
6 subtasks in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
23 shown of 23 papers with code (31 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
14 Jun 2021 2 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedExisting table corpora primarily contain tables extracted from HTML pages, limiting the capability to represent offline database tables.
-
6 May 2021 2 repositories listedExisting work on tabular representation learning jointly models tables and associated text using self-supervised objective functions derived from pretrained language models such as BERT.
-
1 Oct 2019 2 repositories listedThis paper presents the design of our system, namely MTab, for Semantic Web Challenge on Tabular Data to Knowledge Graph Matching (SemTab 2019).
-
25 May 2019 2 repositories listed Syntology ran 0 of 12 samples · 12 unverifiedCorrectly detecting the semantic type of data columns is crucial for data science tasks such as automated data cleaning, schema matching, and data discovery.
-
1 Jan 2025 1 repository listedWe further explore the scenario where training data for the CPA task is available and can be used for selecting demonstrations or fine-tuning the model.
-
27 Jun 2024 1 repository listedWe demonstrate the advantages of statements by applying our model to over 2700 tables from ESG reports.
-
17 Apr 2024 1 repository listedBy leveraging the actual structure and content of tables from Chinese financial announcements, we have developed the first extensive table annotation dataset in this domain.
-
27 Oct 2023 1 repository listed Syntology ran 8 of 12 samples · 4 unverifiedWe introduce ArcheType, a simple, practical method for context sampling, prompt serialization, model querying, and label remapping, which enables large language models to solve CTA problems in a fully zero-shot manner.
-
1 Jun 2023 1 repository listedColumn type annotation is the task of annotating the columns of a relational table with the semantic type of the values contained in each column.
-
27 Mar 2023 1 repository listedTo this end, we propose a new large-scale dataset named Table Recognition Set (TabRecSet) with diverse table forms sourcing from multiple scenarios in the wild, providing complete annotation dedicated to end-to-end TR…
-
10 Jan 2023 1 repository listedIndividual cells and columns are assigned to KG entities and classes to disambiguate their meaning.
-
9 Jan 2023 1 repository listedThis paper presents the WDC Schema.
-
1 Oct 2021 1 repository listedA large portion of structured data does not yet reap the benefits of the Semantic Web.
-
1 Oct 2021 1 repository listedWhile tables are a rich source of structured information, their automated use is oftentimes prevented by the inherent ambiguity contained within.
-
5 Apr 2021 1 repository listed Syntology ran 0 of 3 samples · 3 unverifiedInferring meta information about tables, such as column headers or relationships between columns, is an active research topic in data management as we find many tables are missing some of this information.
-
1 Mar 2021 1 repository listedWe present our publicly available semantic annotator bbw (boosted by wiki) tested at the second Semantic Web Challenge on Tabular Data to Knowledge Graph Matching (SemTab2020).
-
17 Feb 2021 1 repository listedExisting work linearize table cells and heavily rely on modifying deep language models such as BERT which only captures related cells information in the same table.
-
1 Nov 2020 1 repository listedTable annotation is a key task to improve querying the Web and support the Knowledge Graph population from legacy sources (tables).
-
26 Jun 2020 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedIn this paper, we present TURL, a novel framework that introduces the pre-training/fine-tuning paradigm to relational Web tables.
-
30 May 2019 1 repository listedThe usefulness of tabular data such as web tables critically depends on understanding their semantics.
-
4 Nov 2018 1 repository listedAutomatically annotating column types with knowledge base (KB) concepts is a critical task to gain a basic understanding of web tables.
-
1 Jul 2017 1 repository listedUntil now, error type performance for Grammatical Error Correction (GEC) systems could only be measured in terms of recall because system output is not annotated.
-
1 Mar 2017 1 repository listedThis paper contributes to improve the understanding of the utility of different features for web table to knowledge base matching by reimplementing different matching techniques as well as similarity score aggregation…
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections