{"url":"/dataset/tough-tables","name":"Tough Tables","full_name":null,"description_markdown":"The ToughTables (2T) dataset was created for the [SemTab](http://www.cs.ox.ac.uk/isg/challenges/sem-tab/) challenge and includes 180 tables in total. The tables in this dataset can be categorized in two groups: the control (CTRL) group tables and tough (TOUGH) group tables. \r\n\r\nThe CTRL group contains 60 tables generated by querying the DBpedia SPARQL endpoint and tables collected from Wikipedia and their characteristic is that they are easy to annotate. The TOUGH group contains 120 tables mainly scraped from the web, some containing misspelled words and nicknames/homonyms and their characteristic is that they are hard to annotate. In both groups some tables were generated by the authors where they added noise to the collected tables.\r\n\r\nThe dataset was annotated for two tasks using DBpedia (DBP) types and entities and WikiData (WD): Column Type Annotation (CTA) and Cell Entity Annotation (CEA). In the table below the number of columns annotated for the CTA and number of cells annotated for the CEA task as well as the number of classes used are listed.\r\n\r\n|     | Annotations| Classes |\r\n|-----|---------|---------|\r\n| DBP-Column Type Annotation | 540     | 39 |\r\n| DBP-Cell Entity Annotation | 663,656 | 16,023  |\r\n| WD-Column Type Annotation| 540 | 276 |\r\n| WD-Cell Entity Annotation | 667,244 | 24,653 |","description_withheld":null,"homepage":"https://zenodo.org/record/6211551","introduced_date":"2020-11-01","introduced_date_note":null,"introduced_by":{"paper":"/paper/tough-tables-carefully-evaluating-entity","title":"Tough Tables: Carefully Evaluating Entity Linking for Tabular Data","first_author":"Vincenzo Cutrona","url":null},"license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/legalcode"},"modalities":[{"name":"Tabular","url":"/datasets/modality/tabular"}],"tasks":[{"name":"Table annotation","url":"/task/table-annotation","datasets_with_task":"/datasets/task/table-annotation"},{"name":"Column Type Annotation","url":"/task/column-type-annotation","datasets_with_task":"/datasets/task/column-type-annotation"},{"name":"Cell Entity Annotation","url":"/task/cell-entity-annotation","datasets_with_task":"/datasets/task/cell-entity-annotation"}],"languages":[],"variants":["Tough Tables","ToughTables-DBP","ToughTables-WD"],"data_loaders":[],"num_papers_in_archive":11,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/column-type-annotation-on-toughtables-dbp","task":"Column Type Annotation","dataset_variant":"ToughTables-DBP","rows":6,"metrics":["F1 (%)"],"first_row_in_archive_order":{"model":"KGCODE-Tab","paper":"/paper/kgcode-tab-results-for-semtab-2022","metrics":{"F1 (%)":"48"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/cell-entity-annotation-on-toughtables-dbp","task":"Cell Entity Annotation","dataset_variant":"ToughTables-DBP","rows":5,"metrics":["F1 (%)"],"first_row_in_archive_order":{"model":"DAGOBAH","paper":"/paper/dagobah-table-and-graph-contexts-for","metrics":{"F1 (%)":"94.5"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/cell-entity-annotation-on-toughtables-wd","task":"Cell Entity Annotation","dataset_variant":"ToughTables-WD","rows":5,"metrics":["F1 (%)"],"first_row_in_archive_order":{"model":"DAGOBAH","paper":"/paper/from-heuristics-to-language-models-a-journey","metrics":{"F1 (%)":"94.5"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/column-type-annotation-on-toughtables-wd","task":"Column Type Annotation","dataset_variant":"ToughTables-WD","rows":5,"metrics":["F1 (%)"],"first_row_in_archive_order":{"model":"DAGOBAH","paper":"/paper/dagobah-table-and-graph-contexts-for","metrics":{"F1 (%)":"83.2"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/results-of-semtab-2022","title":"Results of SemTab 2022","date":"2022-10-25","rows_on_this_dataset":2,"code_links":0,"syntology":null},{"paper":"/paper/kgcode-tab-results-for-semtab-2022","title":"KGCODE-Tab Results for SemTab 2022","date":"2022-10-25","rows_on_this_dataset":3,"code_links":0,"syntology":null},{"paper":"/paper/from-heuristics-to-language-models-a-journey","title":"From Heuristics to Language Models: A Journey Through the Universe of Semantic Table Interpretation with DAGOBAH","date":"2022-10-25","rows_on_this_dataset":2,"code_links":0,"syntology":null},{"paper":"/paper/magic-mining-an-augmented-graph-using-ink","title":"MAGIC: Mining an Augmented Graph using INK, starting from a CSV","date":"2021-10-01","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/kepler-asi-at-semtab-2021","title":"Kepler-aSI at SemTab 2021","date":"2021-10-01","rows_on_this_dataset":4,"code_links":0,"syntology":null},{"paper":"/paper/jentab-meets-semtab-2021-s-new-challenges","title":"JenTab Meets SemTab 2021's New Challenges","date":"2021-10-01","rows_on_this_dataset":4,"code_links":1,"syntology":null},{"paper":"/paper/dagobah-table-and-graph-contexts-for","title":"DAGOBAH: Table and Graph Contexts for Eﬀicient Semantic Annotation of Tabular Data","date":"2021-10-01","rows_on_this_dataset":4,"code_links":0,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}