Datasets › WDC SOTAB V2
WDC SOTAB V2
SOTAB V2 features two annotation tasks: Column Type Annotation (CTA) and Columns Property Annotation (CPA). The goal of the Column Type Annotation (CTA) task is to annotate the columns of a table using 82 types from the Schema.org vocabulary, such as telephone, Duration, Mass, or Organization. The goal of the Columns Property Annotation (CPA) task is to annotate pairs of table columns with one out of 108 Schema.org properties, such as gtin, startDate, priceValidUntil, or recipeIngredient. The benchmark consists of 45,834 tables annotated for CTA and 30,220 tables annotated for CPA originating from 55,511 different websites. The tables are split into training-, validation- and test sets for both tasks. The tables cover 17 popular Schema.org types including Product, LocalBusiness, Event, and JobPosting.
Some characteristics for the different tasks are provided in the table below, where "Columns" refers to the number of columns/column pairs labeled and "Classes" to the number of unique classes used for annotation.
| Train | Validation | Test | Classes | ||||
|---|---|---|---|---|---|---|---|
| Tables | Columns | Tables | Columns | Tables | Columns | ||
| CTA | 44,769 | 116,887 | 456 | 1,769 | 609 | 1,851 | 82 |
| CPA | 29,158 | 109,994 | 497 | 2,459 | 565 | 2,340 | 108 |
Benchmarks archive 2025-07-28
All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Column Type Annotation | WDC SOTAB V2 | TorchicTab Micro F1 89.66 | TorchicTab: Semantic Table Annotation with Wikidata and... | — | 5 | Compare |
| Columns Property Annotation | WDC SOTAB V2 | TorchicTab Micro F1 87.11 | TorchicTab: Semantic Table Annotation with Wikidata and... | — | 4 | Compare |
Papers archive 2025-07-28
5 shown of 5 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 7. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| TorchicTab: Semantic Table Annotation with Wikidata and Language Models | 0 | 2 | 20 Nov 2023 | not harvested |
| Semantic Annotation of Tabular Data for Machine-to-Machine Interoperability via Neuro-Symbolic Anchoring | 0 | 2 | 20 Nov 2023 | not harvested |
| Exploring Naive Bayes Classifiers for Tabular Data to Knowledge Graph Matching | 0 | 2 | 20 Nov 2023 | not harvested |
| DREIFLUSS: A Minimalist Approach for Table Matching | 0 | 2 | 20 Nov 2023 | not harvested |
| Column Type Annotation using ChatGPT | 1 | 1 | 1 Jun 2023 | not harvested |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- WDC SOTAB V2
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections