Datasets › WikiTables-TURL

WikiTables-TURL

Introduced by Xiang Deng et al. in TURL: Table Understanding through Representation Learning26 Jun 2020 archive 2025-07-28

The WikiTables-TURL dataset was constructed by the authors of TURL and is based on the WikiTable corpus, which is a large collection of Wikipedia tables. The dataset consists of 580,171 tables divided into fixed training, validation and testing splits. Additionally, the dataset contains metadata about each table, such as the table name, table caption and column headers.

406,706 of these tables are annotated for the Column Type Annotation (CTA) task, 55,970 tables for the Columns Property Annotation (CPA) task and 200,744 tables for the Cell Entity Annotation (CEA) task. As classes for the CTA and CPA, Freebase's types and relations were used, whereas for the CEA task entities from Freebase were used. The table below lists the total annotated columns (or cells in the case of CEA) for each split and for each task as well as the number of classes used for annotation.

Training Validation Testing Classes
CTA 628,254 13,391 13,025 255
CPA 62,954 2,175 2,072 121
CEA 1,264,217 76,720 225,777 1,787,737

The authors have made the dataset and its variants publicly available for download.

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

3 shown of 3 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 7. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Watchog: A Light-weight Contrastive Learning based Framework for Column Annotation 0 2 12 Dec 2023 not harvested
Annotating Columns with Pre-trained Language Models 1 2 5 Apr 2021 ran 0 of 3 samples (3 unverified)
TURL: Table Understanding through Representation Learning 1 3 26 Jun 2020 ran 1 of 1 samples (0 unverified)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • WikiTables-TURL
  • WikiTables-TURL-CTA
  • WikiTables-TURL-CPA
  • WikiTables-TURL-CEA

4 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections