{"url":"/dataset/wdc-sotab-v2","name":"WDC SOTAB V2","full_name":null,"description_markdown":"SOTAB V2 features two annotation tasks: Column Type Annotation (CTA) and Columns Property Annotation (CPA). The goal of the Column Type Annotation (CTA) task is to annotate the columns of a table using 82 types from the Schema.org vocabulary, such as telephone, Duration, Mass, or Organization. The goal of the Columns Property Annotation (CPA) task is to annotate pairs of table columns with one out of 108 Schema.org properties, such as gtin, startDate, priceValidUntil, or recipeIngredient. The benchmark consists of 45,834 tables annotated for CTA and 30,220 tables annotated for CPA originating from 55,511 different websites. The tables are split into training-, validation- and test sets for both tasks. The tables cover 17 popular Schema.org types including Product, LocalBusiness, Event, and JobPosting.\r\n\r\nSome characteristics for the different tasks are provided in the table below, where \"Columns\" refers to the number of columns/column pairs labeled and \"Classes\" to the number of unique classes used for annotation.\r\n\r\n|     | Train  |         | Validation |         | Test   |         | Classes |\r\n|-----|--------|---------|------------|---------|--------|---------|---------|\r\n|     | Tables | Columns | Tables     | Columns | Tables | Columns |         |\r\n| CTA | 44,769 | 116,887 | 456        | 1,769   | 609    | 1,851   | 82      |\r\n| CPA | 29,158 | 109,994 | 497        | 2,459   | 565    | 2,340   | 108     |","description_withheld":null,"homepage":"https://webdatacommons.org/structureddata/sotab/v2/","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[{"name":"Tabular","url":"/datasets/modality/tabular"}],"tasks":[{"name":"Data Integration","url":"/task/data-integration","datasets_with_task":"/datasets/task/data-integration"},{"name":"Table annotation","url":"/task/table-annotation","datasets_with_task":"/datasets/task/table-annotation"},{"name":"Column Type Annotation","url":"/task/column-type-annotation","datasets_with_task":"/datasets/task/column-type-annotation"},{"name":"Columns Property Annotation","url":"/task/columns-property-annotation","datasets_with_task":"/datasets/task/columns-property-annotation"}],"languages":[],"variants":["WDC SOTAB V2"],"data_loaders":[],"num_papers_in_archive":7,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/column-type-annotation-on-wdc-sotab-v2","task":"Column Type Annotation","dataset_variant":"WDC SOTAB V2","rows":5,"metrics":["Micro F1"],"first_row_in_archive_order":{"model":"TorchicTab","paper":"/paper/torchictab-semantic-table-annotation-with","metrics":{"Micro F1":"89.66"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/columns-property-annotation-on-wdc-sotab-v2","task":"Columns Property Annotation","dataset_variant":"WDC SOTAB V2","rows":4,"metrics":["Micro F1"],"first_row_in_archive_order":{"model":"TorchicTab","paper":"/paper/torchictab-semantic-table-annotation-with","metrics":{"Micro F1":"87.11"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/torchictab-semantic-table-annotation-with","title":"TorchicTab: Semantic Table Annotation with Wikidata and Language Models","date":"2023-11-20","rows_on_this_dataset":2,"code_links":0,"syntology":null},{"paper":"/paper/semantic-annotation-of-tabular-data-for","title":"Semantic Annotation of Tabular Data for Machine-to-Machine Interoperability via Neuro-Symbolic Anchoring","date":"2023-11-20","rows_on_this_dataset":2,"code_links":0,"syntology":null},{"paper":"/paper/exploring-naive-bayes-classifiers-for-tabular","title":"Exploring Naive Bayes Classifiers for Tabular Data to Knowledge Graph Matching","date":"2023-11-20","rows_on_this_dataset":2,"code_links":0,"syntology":null},{"paper":"/paper/dreifluss-a-minimalist-approach-for-table","title":"DREIFLUSS: A Minimalist Approach for Table Matching","date":"2023-11-20","rows_on_this_dataset":2,"code_links":0,"syntology":null},{"paper":"/paper/column-type-annotation-using-chatgpt","title":"Column Type Annotation using ChatGPT","date":"2023-06-01","rows_on_this_dataset":1,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}