{"url":"/task/table-annotation","name":"Table annotation","slug":"table-annotation","description_markdown":"**Table annotation** is the task of annotating a table with terms/concepts from knowledge graph or database schema. Table annotation is typically broken down into the following five subtasks: \r\n\r\n1. Cell Entity Annotation ([CEA](https://paperswithcode.com/task/cell-entity-annotation))\r\n2. Column Type Annotation ([CTA](https://paperswithcode.com/task/column-type-annotation))\r\n3. Column Property Annotation ([CPA](https://paperswithcode.com/task/columns-property-annotation))\r\n4. [Table Type Detection](https://paperswithcode.com/task/table-type-detection)\r\n5. [Row Annotation](https://paperswithcode.com/task/row-annotation)\r\n\r\nThe [SemTab](http://www.cs.ox.ac.uk/isg/challenges/sem-tab/) challenge is closely related to the Table Annotation problem. It is a yearly challenge which focuses on the first three tasks of table annotation and its purpose is to benchmark different table annotation systems.","categories":[{"name":"Knowledge Base","url":"/area/knowledge-base"},{"name":"Natural Language Processing","url":"/area/natural-language-processing"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":31,"papers_with_code":23,"benchmarks":0,"benchmark_tables_in_archive":0,"benchmark_tables_shown":0,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":10,"subtasks":6,"parent_tasks":1},"benchmarks":[],"datasets":[{"url":"/dataset/gittables","name":"GitTables","full_name":"","num_papers_in_archive":16},{"url":"/dataset/t2dv2","name":"T2Dv2","full_name":"","num_papers_in_archive":14},{"url":"/dataset/tough-tables","name":"Tough Tables","full_name":"","num_papers_in_archive":11},{"url":"/dataset/wdc-sotab-v2","name":"WDC SOTAB V2","full_name":"","num_papers_in_archive":7},{"url":"/dataset/wikitables-turl","name":"WikiTables-TURL","full_name":"","num_papers_in_archive":7},{"url":"/dataset/viznet-sato","name":"VizNet-Sato","full_name":"","num_papers_in_archive":4},{"url":"/dataset/wikipediags","name":"WikipediaGS","full_name":"","num_papers_in_archive":4},{"url":"/dataset/information-extraction-from-tables","name":"Information Extraction from Tables","full_name":"Extraction materials compositions from tables of materials science research papers","num_papers_in_archive":3},{"url":"/dataset/tncr-dataset","name":"TNCR Dataset","full_name":"Table Net Detection and Classification Dataset","num_papers_in_archive":2},{"url":"/dataset/wdc-sotab","name":"WDC SOTAB","full_name":"","num_papers_in_archive":2}],"subtasks":[{"url":"/task/cell-entity-annotation","name":"Cell Entity Annotation"},{"url":"/task/column-type-annotation","name":"Column Type Annotation"},{"url":"/task/columns-property-annotation","name":"Columns Property Annotation"},{"url":"/task/metric-type-identification","name":"Metric-Type Identification"},{"url":"/task/row-annotation","name":"Row Annotation"},{"url":"/task/table-type-detection","name":"Table Type Detection"}],"parent_tasks":[{"url":"/task/data-integration","name":"Data Integration"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":23,"of":23,"tagged_in_all":31,"items":[{"url":"/paper/gittables-a-large-scale-corpus-of-relational","title":"GitTables: A Large-Scale Corpus of Relational Tables","date":"2021-06-14","arxiv_id":"2106.07258","repositories_listed":2,"syntology":{"n":1,"n_ran":0,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/tabbie-pretrained-representations-of-tabular","title":"TABBIE: Pretrained Representations of Tabular Data","date":"2021-05-06","arxiv_id":"2105.02584","repositories_listed":2,"syntology":null},{"url":"/paper/mtab-matching-tabular-data-to-knowledge-graph","title":"MTab: Matching Tabular Data to Knowledge Graph using Probability Models","date":"2019-10-01","arxiv_id":"1910.00246","repositories_listed":2,"syntology":null},{"url":"/paper/sherlock-a-deep-learning-approach-to-semantic","title":"Sherlock: A Deep Learning Approach to Semantic Data Type Detection","date":"2019-05-25","arxiv_id":"1905.10688","repositories_listed":2,"syntology":{"n":12,"n_ran":0,"n_unverified":12,"n_pointer_only":0}},{"url":"/paper/column-property-annotation-using-large","title":"Column Property Annotation using Large Language Models","date":"2025-01-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/statements-universal-information-extraction","title":"Statements: Universal Information Extraction from Tables with Large Language Models for ESG KPIs","date":"2024-06-27","arxiv_id":"2406.19102","repositories_listed":1,"syntology":null},{"url":"/paper/synthesizing-realistic-data-for-table","title":"Synthesizing Realistic Data for Table Recognition","date":"2024-04-17","arxiv_id":"2404.11100","repositories_listed":1,"syntology":null},{"url":"/paper/archetype-a-novel-framework-for-open-source","title":"ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models","date":"2023-10-27","arxiv_id":"2310.18208","repositories_listed":1,"syntology":{"n":12,"n_ran":8,"n_unverified":4,"n_pointer_only":0}},{"url":"/paper/column-type-annotation-using-chatgpt","title":"Column Type Annotation using ChatGPT","date":"2023-06-01","arxiv_id":"2306.00745","repositories_listed":1,"syntology":null},{"url":"/paper/a-large-scale-dataset-for-end-to-end-table","title":"A large-scale dataset for end-to-end table recognition in the wild","date":"2023-03-27","arxiv_id":"2303.14884","repositories_listed":1,"syntology":null},{"url":"/paper/biodivtab-semantic-table-annotation-benchmark","title":"BiodivTab: Semantic Table Annotation Benchmark Construction, Analysis, and New Additions","date":"2023-01-10","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/sotab-the-wdc-schema-org-table-annotation","title":"SOTAB: The WDC Schema.org Table Annotation Benchmark","date":"2023-01-09","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/magic-mining-an-augmented-graph-using-ink","title":"MAGIC: Mining an Augmented Graph using INK, starting from a CSV","date":"2021-10-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/jentab-meets-semtab-2021-s-new-challenges","title":"JenTab Meets SemTab 2021's New Challenges","date":"2021-10-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/annotating-columns-with-pre-trained-language","title":"Annotating Columns with Pre-trained Language Models","date":"2021-04-05","arxiv_id":"2104.01785","repositories_listed":1,"syntology":{"n":3,"n_ran":0,"n_unverified":3,"n_pointer_only":0}},{"url":"/paper/bbw-matching-csv-to-wikidata-via-meta-lookup","title":"bbw: Matching CSV to Wikidata via Meta-lookup","date":"2021-03-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/tcn-table-convolutional-network-for-web-table","title":"TCN: Table Convolutional Network for Web Table Interpretation","date":"2021-02-17","arxiv_id":"2102.09460","repositories_listed":1,"syntology":null},{"url":"/paper/tough-tables-carefully-evaluating-entity","title":"Tough Tables: Carefully Evaluating Entity Linking for Tabular Data","date":"2020-11-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/turl-table-understanding-through","title":"TURL: Table Understanding through Representation Learning","date":"2020-06-26","arxiv_id":"2006.14806","repositories_listed":1,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/190600781","title":"Learning Semantic Annotations for Tabular Data","date":"2019-05-30","arxiv_id":"1906.00781","repositories_listed":1,"syntology":null},{"url":"/paper/colnet-embedding-the-semantics-of-web-tables","title":"ColNet: Embedding the Semantics of Web Tables for Column Type Prediction","date":"2018-11-04","arxiv_id":"1811.01304","repositories_listed":1,"syntology":null},{"url":"/paper/automatic-annotation-and-evaluation-of-error","title":"Automatic Annotation and Evaluation of Error Types for Grammatical Error Correction","date":"2017-07-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/matching-web-tables-to-dbpedia-a-feature","title":"Matching web tables to DBpedia-A feature utility study","date":"2017-03-01","arxiv_id":null,"repositories_listed":1,"syntology":null}],"syntology_records":5,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":5,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}