{"url":"/task/column-type-annotation","name":"Column Type Annotation","slug":"column-type-annotation","description_markdown":"**Column type annotation** (CTA) refers to the task of predicting the semantic type of a table column and is a subtask of [Table Annotation](https://paperswithcode.com/task/table-annotation). The labels that are usually used in a CTA problem are semantic types from vocabularies like DBpedia, Schema.org or WikiData. Some examples are: *Book*, *Country*, *LocalBusiness* etc.\r\n\r\nCTA can be either treated as a multi-class classification problem where a column is annotated by only one semantic type or as multi-label classification problem where a column can be annotated using multiple semantic types.","categories":[{"name":"Knowledge Base","url":"/area/knowledge-base"},{"name":"Natural Language Processing","url":"/area/natural-language-processing"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":34,"papers_with_code":19,"benchmarks":12,"benchmark_tables_in_archive":12,"benchmark_tables_shown":12,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":10,"subtasks":0,"parent_tasks":1},"benchmarks":[{"leaderboard":"/sota/column-type-annotation-on-biodivtab","slug":"column-type-annotation-on-biodivtab","dataset":"BiodivTab","dataset_url":"/dataset/biodivtab","rows_in_archive":6,"metrics":["F1 (%)"],"first_row_in_archive_order":{"model":"KGCODE-Tab","paper_title":"KGCODE-Tab Results for SemTab 2022","paper_url":"/paper/kgcode-tab-results-for-semtab-2022","paper_date":"2022-10-25","arxiv_id":null,"code_links":[],"syntology":null}},{"leaderboard":"/sota/column-type-annotation-on-toughtables-dbp","slug":"column-type-annotation-on-toughtables-dbp","dataset":"ToughTables-DBP","dataset_url":"/dataset/tough-tables","rows_in_archive":6,"metrics":["F1 (%)"],"first_row_in_archive_order":{"model":"KGCODE-Tab","paper_title":"KGCODE-Tab Results for SemTab 2022","paper_url":"/paper/kgcode-tab-results-for-semtab-2022","paper_date":"2022-10-25","arxiv_id":null,"code_links":[],"syntology":null}},{"leaderboard":"/sota/column-type-annotation-on-toughtables-wd","slug":"column-type-annotation-on-toughtables-wd","dataset":"ToughTables-WD","dataset_url":"/dataset/tough-tables","rows_in_archive":5,"metrics":["F1 (%)"],"first_row_in_archive_order":{"model":"DAGOBAH","paper_title":"DAGOBAH: Table and Graph Contexts for Eﬀicient Semantic Annotation of Tabular Data","paper_url":"/paper/dagobah-table-and-graph-contexts-for","paper_date":"2021-10-01","arxiv_id":null,"code_links":[],"syntology":null}},{"leaderboard":"/sota/column-type-annotation-on-wdc-sotab-v2","slug":"column-type-annotation-on-wdc-sotab-v2","dataset":"WDC SOTAB V2","dataset_url":"/dataset/wdc-sotab-v2","rows_in_archive":5,"metrics":["Micro F1"],"first_row_in_archive_order":{"model":"TorchicTab","paper_title":"TorchicTab: Semantic Table Annotation with Wikidata and Language Models","paper_url":"/paper/torchictab-semantic-table-annotation-with","paper_date":"2023-11-20","arxiv_id":null,"code_links":[],"syntology":null}},{"leaderboard":"/sota/column-type-annotation-on-viznet-sato-full","slug":"column-type-annotation-on-viznet-sato-full","dataset":"VizNet-Sato-Full","dataset_url":"/dataset/viznet-sato","rows_in_archive":4,"metrics":["Macro-F1","Weighted-F1"],"first_row_in_archive_order":{"model":"Watchog","paper_title":"Watchog: A Light-weight Contrastive Learning based Framework for Column Annotation","paper_url":"/paper/watchog-a-light-weight-contrastive-learning","paper_date":"2023-12-12","arxiv_id":null,"code_links":[],"syntology":null}},{"leaderboard":"/sota/column-type-annotation-on-gittables-semtab","slug":"column-type-annotation-on-gittables-semtab","dataset":"GitTables-SemTab-DBP","dataset_url":"/dataset/gittables-semtab","rows_in_archive":3,"metrics":["F1 (%)"],"first_row_in_archive_order":{"model":"KGCODE-Tab","paper_title":"KGCODE-Tab Results for SemTab 2022","paper_url":"/paper/kgcode-tab-results-for-semtab-2022","paper_date":"2022-10-25","arxiv_id":null,"code_links":[],"syntology":null}},{"leaderboard":"/sota/column-type-annotation-on-t2dv2","slug":"column-type-annotation-on-t2dv2","dataset":"T2Dv2","dataset_url":"/dataset/t2dv2","rows_in_archive":3,"metrics":["Accuracy (%)","F1 (%)"],"first_row_in_archive_order":{"model":"HNN + P2Vec","paper_title":"Learning Semantic Annotations for Tabular Data","paper_url":"/paper/190600781","paper_date":"2019-05-30","arxiv_id":"1906.00781","code_links":[{"title":"alan-turing-institute/SemAIDA","url":"https://github.com/alan-turing-institute/SemAIDA"}],"syntology":null}},{"leaderboard":"/sota/column-type-annotation-on-wdc-sotab","slug":"column-type-annotation-on-wdc-sotab","dataset":"WDC SOTAB","dataset_url":"/dataset/wdc-sotab","rows_in_archive":3,"metrics":["Micro F1","Weighted F1"],"first_row_in_archive_order":{"model":"DODUO","paper_title":"SOTAB: The WDC Schema.org Table Annotation Benchmark","paper_url":"/paper/sotab-the-wdc-schema-org-table-annotation","paper_date":"2023-01-09","arxiv_id":null,"code_links":[{"title":"wbsg-uni-mannheim/wdc-sotab","url":"https://github.com/wbsg-uni-mannheim/wdc-sotab"}],"syntology":null}},{"leaderboard":"/sota/column-type-annotation-on-wikitables-turl-cta","slug":"column-type-annotation-on-wikitables-turl-cta","dataset":"WikiTables-TURL-CTA","dataset_url":"/dataset/wikitables-turl","rows_in_archive":3,"metrics":["F1 (%)","Macro-F1"],"first_row_in_archive_order":{"model":"TURL","paper_title":"TURL: Table Understanding through Representation Learning","paper_url":"/paper/turl-table-understanding-through","paper_date":"2020-06-26","arxiv_id":"2006.14806","code_links":[{"title":"sunlab-osu/TURL","url":"https://github.com/sunlab-osu/TURL"}],"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":0}}},{"leaderboard":"/sota/column-type-annotation-on-gittables-semtab-1","slug":"column-type-annotation-on-gittables-semtab-1","dataset":"GitTables-SemTab-SCH","dataset_url":"/dataset/gittables-semtab","rows_in_archive":2,"metrics":["F1 (%)"],"first_row_in_archive_order":{"model":"KGCODE-Tab","paper_title":"KGCODE-Tab Results for SemTab 2022","paper_url":"/paper/kgcode-tab-results-for-semtab-2022","paper_date":"2022-10-25","arxiv_id":null,"code_links":[],"syntology":null}},{"leaderboard":"/sota/column-type-annotation-on-viznet-sato-1","slug":"column-type-annotation-on-viznet-sato-1","dataset":"VizNet-Sato-MultiColumn","dataset_url":"/dataset/viznet-sato","rows_in_archive":2,"metrics":["Macro-F1","Weighted-F1"],"first_row_in_archive_order":{"model":"DODUO","paper_title":"Annotating Columns with Pre-trained Language Models","paper_url":"/paper/annotating-columns-with-pre-trained-language","paper_date":"2021-04-05","arxiv_id":"2104.01785","code_links":[{"title":"megagonlabs/doduo","url":"https://github.com/megagonlabs/doduo"}],"syntology":{"n":3,"n_ran":0,"n_unverified":3,"n_pointer_only":0}}},{"leaderboard":"/sota/column-type-annotation-on-wikipediags-cta","slug":"column-type-annotation-on-wikipediags-cta","dataset":"WikipediaGS-CTA","dataset_url":"/dataset/wikipediags","rows_in_archive":2,"metrics":["Accuracy (%)"],"first_row_in_archive_order":{"model":"TURL","paper_title":"TURL: Table Understanding through Representation Learning","paper_url":"/paper/turl-table-understanding-through","paper_date":"2020-06-26","arxiv_id":"2006.14806","code_links":[{"title":"sunlab-osu/TURL","url":"https://github.com/sunlab-osu/TURL"}],"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":0}}}],"datasets":[{"url":"/dataset/gittables","name":"GitTables","full_name":"","num_papers_in_archive":16},{"url":"/dataset/t2dv2","name":"T2Dv2","full_name":"","num_papers_in_archive":14},{"url":"/dataset/tough-tables","name":"Tough Tables","full_name":"","num_papers_in_archive":11},{"url":"/dataset/biodivtab","name":"BiodivTab","full_name":"","num_papers_in_archive":8},{"url":"/dataset/wdc-sotab-v2","name":"WDC SOTAB V2","full_name":"","num_papers_in_archive":7},{"url":"/dataset/wikitables-turl","name":"WikiTables-TURL","full_name":"","num_papers_in_archive":7},{"url":"/dataset/viznet-sato","name":"VizNet-Sato","full_name":"","num_papers_in_archive":4},{"url":"/dataset/wikipediags","name":"WikipediaGS","full_name":"","num_papers_in_archive":4},{"url":"/dataset/gittables-semtab","name":"GitTables-SemTab","full_name":"","num_papers_in_archive":3},{"url":"/dataset/wdc-sotab","name":"WDC SOTAB","full_name":"","num_papers_in_archive":2}],"subtasks":[],"parent_tasks":[{"url":"/task/table-annotation","name":"Table annotation"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":19,"of":19,"tagged_in_all":34,"items":[{"url":"/paper/tabbie-pretrained-representations-of-tabular","title":"TABBIE: Pretrained Representations of Tabular Data","date":"2021-05-06","arxiv_id":"2105.02584","repositories_listed":2,"syntology":null},{"url":"/paper/sherlock-a-deep-learning-approach-to-semantic","title":"Sherlock: A Deep Learning Approach to Semantic Data Type Detection","date":"2019-05-25","arxiv_id":"1905.10688","repositories_listed":2,"syntology":{"n":12,"n_ran":0,"n_unverified":12,"n_pointer_only":0}},{"url":"/paper/evaluating-knowledge-generation-and-self","title":"Evaluating Knowledge Generation and Self-Refinement Strategies for LLM-based Column Type Annotation","date":"2025-03-04","arxiv_id":"2503.02718","repositories_listed":1,"syntology":null},{"url":"/paper/column-property-annotation-using-large","title":"Column Property Annotation using Large Language Models","date":"2025-01-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/accio-table-understanding-enhanced-via","title":"ACCIO: Table Understanding Enhanced via Contrastive Learning with Aggregations","date":"2024-11-07","arxiv_id":"2411.04443","repositories_listed":1,"syntology":null},{"url":"/paper/kglink-a-column-type-annotation-method-that","title":"KGLink: A column type annotation method that combines knowledge graph and pre-trained language model","date":"2024-06-01","arxiv_id":"2406.00318","repositories_listed":1,"syntology":null},{"url":"/paper/archetype-a-novel-framework-for-open-source","title":"ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models","date":"2023-10-27","arxiv_id":"2310.18208","repositories_listed":1,"syntology":{"n":12,"n_ran":8,"n_unverified":4,"n_pointer_only":0}},{"url":"/paper/chorus-foundation-models-for-unified-data","title":"CHORUS: Foundation Models for Unified Data Discovery and Exploration","date":"2023-06-16","arxiv_id":"2306.09610","repositories_listed":1,"syntology":null},{"url":"/paper/column-type-annotation-using-chatgpt","title":"Column Type Annotation using ChatGPT","date":"2023-06-01","arxiv_id":"2306.00745","repositories_listed":1,"syntology":null},{"url":"/paper/sotab-the-wdc-schema-org-table-annotation","title":"SOTAB: The WDC Schema.org Table Annotation Benchmark","date":"2023-01-09","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/towards-an-approach-based-on-knowledge-graph","title":"Towards an Approach based on Knowledge Graph Refinement for Tabular Data to Knowledge Graph Matching","date":"2022-10-25","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/magic-mining-an-augmented-graph-using-ink","title":"MAGIC: Mining an Augmented Graph using INK, starting from a CSV","date":"2021-10-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/jentab-meets-semtab-2021-s-new-challenges","title":"JenTab Meets SemTab 2021's New Challenges","date":"2021-10-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/annotating-columns-with-pre-trained-language","title":"Annotating Columns with Pre-trained Language Models","date":"2021-04-05","arxiv_id":"2104.01785","repositories_listed":1,"syntology":{"n":3,"n_ran":0,"n_unverified":3,"n_pointer_only":0}},{"url":"/paper/tough-tables-carefully-evaluating-entity","title":"Tough Tables: Carefully Evaluating Entity Linking for Tabular Data","date":"2020-11-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/turl-table-understanding-through","title":"TURL: Table Understanding through Representation Learning","date":"2020-06-26","arxiv_id":"2006.14806","repositories_listed":1,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/sato-contextual-semantic-type-detection-in","title":"Sato: Contextual Semantic Type Detection in Tables","date":"2019-11-14","arxiv_id":"1911.06311","repositories_listed":1,"syntology":{"n":10,"n_ran":1,"n_unverified":9,"n_pointer_only":0}},{"url":"/paper/190600781","title":"Learning Semantic Annotations for Tabular Data","date":"2019-05-30","arxiv_id":"1906.00781","repositories_listed":1,"syntology":null},{"url":"/paper/colnet-embedding-the-semantics-of-web-tables","title":"ColNet: Embedding the Semantics of Web Tables for Column Type Prediction","date":"2018-11-04","arxiv_id":"1811.01304","repositories_listed":1,"syntology":null}],"syntology_records":5,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":1,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}