Datasets › Tough Tables
Tough Tables
The ToughTables (2T) dataset was created for the SemTab challenge and includes 180 tables in total. The tables in this dataset can be categorized in two groups: the control (CTRL) group tables and tough (TOUGH) group tables.
The CTRL group contains 60 tables generated by querying the DBpedia SPARQL endpoint and tables collected from Wikipedia and their characteristic is that they are easy to annotate. The TOUGH group contains 120 tables mainly scraped from the web, some containing misspelled words and nicknames/homonyms and their characteristic is that they are hard to annotate. In both groups some tables were generated by the authors where they added noise to the collected tables.
The dataset was annotated for two tasks using DBpedia (DBP) types and entities and WikiData (WD): Column Type Annotation (CTA) and Cell Entity Annotation (CEA). In the table below the number of columns annotated for the CTA and number of cells annotated for the CEA task as well as the number of classes used are listed.
| Annotations | Classes | |
|---|---|---|
| DBP-Column Type Annotation | 540 | 39 |
| DBP-Cell Entity Annotation | 663,656 | 16,023 |
| WD-Column Type Annotation | 540 | 276 |
| WD-Cell Entity Annotation | 667,244 | 24,653 |
Benchmarks archive 2025-07-28
All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Column Type Annotation | ToughTables-DBP | KGCODE-Tab F1 (%) 48 | KGCODE-Tab Results for SemTab 2022 | — | 6 | Compare |
| Cell Entity Annotation | ToughTables-DBP | DAGOBAH F1 (%) 94.5 | DAGOBAH: Table and Graph Contexts for Efficient Semantic... | — | 5 | Compare |
| Cell Entity Annotation | ToughTables-WD | DAGOBAH F1 (%) 94.5 | From Heuristics to Language Models: A Journey Through... | — | 5 | Compare |
| Column Type Annotation | ToughTables-WD | DAGOBAH F1 (%) 83.2 | DAGOBAH: Table and Graph Contexts for Efficient Semantic... | — | 5 | Compare |
Papers archive 2025-07-28
7 shown of 7 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 11. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Results of SemTab 2022 | 0 | 2 | 25 Oct 2022 | not harvested |
| KGCODE-Tab Results for SemTab 2022 | 0 | 3 | 25 Oct 2022 | not harvested |
| From Heuristics to Language Models: A Journey Through the Universe of Semantic Table Interpretation with DAGOBAH | 0 | 2 | 25 Oct 2022 | not harvested |
| MAGIC: Mining an Augmented Graph using INK, starting from a CSV | 1 | 2 | 1 Oct 2021 | not harvested |
| Kepler-aSI at SemTab 2021 | 0 | 4 | 1 Oct 2021 | not harvested |
| JenTab Meets SemTab 2021's New Challenges | 1 | 4 | 1 Oct 2021 | not harvested |
| DAGOBAH: Table and Graph Contexts for Efficient Semantic Annotation of Tabular Data | 0 | 4 | 1 Oct 2021 | not harvested |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- Tough Tables
- ToughTables-DBP
- ToughTables-WD
3 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections