Browse State-of-the-Art › Column Type Annotation
Column Type Annotation
19 papers with code · 12 benchmarks · 10 datasets archive 2025-07-28
Column type annotation (CTA) refers to the task of predicting the semantic type of a table column and is a subtask of Table Annotation. The labels that are usually used in a CTA problem are semantic types from vocabularies like DBpedia, Schema.org or WikiData. Some examples are: Book, Country, LocalBusiness etc.
CTA can be either treated as a multi-class classification problem where a column is annotated by only one semantic type or as multi-label classification problem where a column can be annotated using multiple semantic types.
Description from the archive archive 2025-07-28; Papers-with-Code links inside it are rewritten to this site.
Benchmarks archive 2025-07-28
12 leaderboard tables shown for this task, 12 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 12 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
10 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
19 shown of 19 papers with code (34 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
6 May 2021 2 repositories listedExisting work on tabular representation learning jointly models tables and associated text using self-supervised objective functions derived from pretrained language models such as BERT.
-
25 May 2019 2 repositories listed Syntology ran 0 of 12 samples · 12 unverifiedCorrectly detecting the semantic type of data columns is crucial for data science tasks such as automated data cleaning, schema matching, and data discovery.
-
4 Mar 2025 1 repository listedThe strategies include using LLMs to generate term definitions, error-based refinement of term definitions, self-correction, and fine-tuning using examples and term definitions.
-
1 Jan 2025 1 repository listedWe further explore the scenario where training data for the CPA task is available and can be used for selecting demonstrations or fine-tuning the model.
-
7 Nov 2024 1 repository listedWith this idea, in this paper, we introduce ACCIO, tAble understanding enhanCed via Contrastive learnIng with aggregatiOns, a novel approach to enhancing table understanding by contrasting original tables with their…
-
KGLink: A column type annotation method that combines knowledge graph and pre-trained language model1 Jun 2024 1 repository listedBy leveraging the strengths of KGLink, we successfully surmount challenges related to type granularity and valuable context issues, establishing it as a robust solution for the semantic annotation of tabular data.
-
27 Oct 2023 1 repository listed Syntology ran 8 of 12 samples · 4 unverifiedWe introduce ArcheType, a simple, practical method for context sampling, prompt serialization, model querying, and label remapping, which enables large language models to solve CTA problems in a fully zero-shot manner.
-
16 Jun 2023 1 repository listedOn all three tasks, we show that a foundation-model-based approach outperforms the task-specific models and so the state of the art.
-
1 Jun 2023 1 repository listedColumn type annotation is the task of annotating the columns of a relational table with the semantic type of the values contained in each column.
-
9 Jan 2023 1 repository listedThis paper presents the WDC Schema.
-
Towards an Approach based on Knowledge Graph Refinement for Tabular Data to Knowledge Graph Matching25 Oct 2022 1 repository listedThis paper presents our contribution to the Accuracy Track of Semantic Web Challenge on Tabular Data to Knowledge Graph Matching (SemTab).
-
1 Oct 2021 1 repository listedA large portion of structured data does not yet reap the benefits of the Semantic Web.
-
1 Oct 2021 1 repository listedWhile tables are a rich source of structured information, their automated use is oftentimes prevented by the inherent ambiguity contained within.
-
5 Apr 2021 1 repository listed Syntology ran 0 of 3 samples · 3 unverifiedInferring meta information about tables, such as column headers or relationships between columns, is an active research topic in data management as we find many tables are missing some of this information.
-
1 Nov 2020 1 repository listedTable annotation is a key task to improve querying the Web and support the Knowledge Graph population from legacy sources (tables).
-
26 Jun 2020 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedIn this paper, we present TURL, a novel framework that introduces the pre-training/fine-tuning paradigm to relational Web tables.
-
14 Nov 2019 1 repository listed Syntology ran 1 of 10 samples · 9 unverifiedDetecting the semantic types of data columns in relational tables is important for various data preparation and information retrieval tasks such as data cleaning, schema matching, data discovery, and semantic search.
-
30 May 2019 1 repository listedThe usefulness of tabular data such as web tables critically depends on understanding their semantics.
-
4 Nov 2018 1 repository listedAutomatically annotating column types with knowledge base (KB) concepts is a critical task to gain a basic understanding of web tables.
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections