Datasets › WikiTables-TURL
WikiTables-TURL
The WikiTables-TURL dataset was constructed by the authors of TURL and is based on the WikiTable corpus, which is a large collection of Wikipedia tables. The dataset consists of 580,171 tables divided into fixed training, validation and testing splits. Additionally, the dataset contains metadata about each table, such as the table name, table caption and column headers.
406,706 of these tables are annotated for the Column Type Annotation (CTA) task, 55,970 tables for the Columns Property Annotation (CPA) task and 200,744 tables for the Cell Entity Annotation (CEA) task. As classes for the CTA and CPA, Freebase's types and relations were used, whereas for the CEA task entities from Freebase were used. The table below lists the total annotated columns (or cells in the case of CEA) for each split and for each task as well as the number of classes used for annotation.
| Training | Validation | Testing | Classes | |
|---|---|---|---|---|
| CTA | 628,254 | 13,391 | 13,025 | 255 |
| CPA | 62,954 | 2,175 | 2,072 | 121 |
| CEA | 1,264,217 | 76,720 | 225,777 | 1,787,737 |
The authors have made the dataset and its variants publicly available for download.
Benchmarks archive 2025-07-28
All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Column Type Annotation | WikiTables-TURL-CTA | TURL F1 (%) 94.75 | TURL: Table Understanding through Representation Learning | sunlab-osu/TURL | 3 | Compare |
| Columns Property Annotation | WikiTables-TURL-CPA | TURL F1 (%) 94.91 | TURL: Table Understanding through Representation Learning | sunlab-osu/TURL | 3 | Compare |
| Cell Entity Annotation | WikiTables-TURL-CEA | TURL F1 (%) 68 | TURL: Table Understanding through Representation Learning | sunlab-osu/TURL | 1 | Compare |
Papers archive 2025-07-28
3 shown of 3 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 7. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Watchog: A Light-weight Contrastive Learning based Framework for Column Annotation | 0 | 2 | 12 Dec 2023 | not harvested |
| Annotating Columns with Pre-trained Language Models | 1 | 2 | 5 Apr 2021 | ran 0 of 3 samples (3 unverified) |
| TURL: Table Understanding through Representation Learning | 1 | 3 | 26 Jun 2020 | ran 1 of 1 samples (0 unverified) |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- WikiTables-TURL
- WikiTables-TURL-CTA
- WikiTables-TURL-CPA
- WikiTables-TURL-CEA
4 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections