Datasets › Tough Tables

Tough Tables

Introduced by Vincenzo Cutrona et al. in Tough Tables: Carefully Evaluating Entity Linking for Tabular Data1 Nov 2020 archive 2025-07-28

The ToughTables (2T) dataset was created for the SemTab challenge and includes 180 tables in total. The tables in this dataset can be categorized in two groups: the control (CTRL) group tables and tough (TOUGH) group tables.

The CTRL group contains 60 tables generated by querying the DBpedia SPARQL endpoint and tables collected from Wikipedia and their characteristic is that they are easy to annotate. The TOUGH group contains 120 tables mainly scraped from the web, some containing misspelled words and nicknames/homonyms and their characteristic is that they are hard to annotate. In both groups some tables were generated by the authors where they added noise to the collected tables.

The dataset was annotated for two tasks using DBpedia (DBP) types and entities and WikiData (WD): Column Type Annotation (CTA) and Cell Entity Annotation (CEA). In the table below the number of columns annotated for the CTA and number of cells annotated for the CEA task as well as the number of classes used are listed.

Annotations Classes
DBP-Column Type Annotation 540 39
DBP-Cell Entity Annotation 663,656 16,023
WD-Column Type Annotation 540 276
WD-Cell Entity Annotation 667,244 24,653

Benchmarks archive 2025-07-28

All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Column Type Annotation ToughTables-DBP KGCODE-Tab F1 (%) 48 KGCODE-Tab Results for SemTab 2022 — 6 Compare
Cell Entity Annotation ToughTables-DBP DAGOBAH F1 (%) 94.5 DAGOBAH: Table and Graph Contexts for Efficient Semantic... — 5 Compare
Cell Entity Annotation ToughTables-WD DAGOBAH F1 (%) 94.5 From Heuristics to Language Models: A Journey Through... — 5 Compare
Column Type Annotation ToughTables-WD DAGOBAH F1 (%) 83.2 DAGOBAH: Table and Graph Contexts for Efficient Semantic... — 5 Compare

Papers archive 2025-07-28

7 shown of 7 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 11. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Results of SemTab 2022 0 2 25 Oct 2022 not harvested
KGCODE-Tab Results for SemTab 2022 0 3 25 Oct 2022 not harvested
From Heuristics to Language Models: A Journey Through the Universe of Semantic Table Interpretation with DAGOBAH 0 2 25 Oct 2022 not harvested
MAGIC: Mining an Augmented Graph using INK, starting from a CSV 1 2 1 Oct 2021 not harvested
Kepler-aSI at SemTab 2021 0 4 1 Oct 2021 not harvested
JenTab Meets SemTab 2021's New Challenges 1 4 1 Oct 2021 not harvested
DAGOBAH: Table and Graph Contexts for Efficient Semantic Annotation of Tabular Data 0 4 1 Oct 2021 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • Tough Tables
  • ToughTables-DBP
  • ToughTables-WD

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections