{"url":"/dataset/wdc-products","name":"WDC LSPM","full_name":null,"description_markdown":"Many e-shops have started to mark-up product data within their HTML pages using the schema.org vocabulary. The Web Data Commons project regularly extracts such data from the Common Crawl, a large public web crawl. The Web Data Commons Training and Test Sets for Large-Scale Product Matching contain product offers from different e-shops in the form of binary product pairs (with corresponding label \"match\" or \"no match\") for four product categories, computers, cameras, watches and shoes. \r\n\r\nIn order to support the evaluation of machine learning-based matching methods, the data is split into training, validation and test sets. For each product category, we provide training sets in four different sizes (2.000-70.000 pairs). Furthermore there are sets of ids for each training set for a possible validation split (stratified random draw) available. The test set for each product category consists of 1.100 product pairs. The labels of the test sets were manually checked while those of the training sets were derived using shared product identifiers from the Web via weak supervision. \r\n\r\nThe data stems from the WDC Product Data Corpus for Large-Scale Product Matching - Version 2.0 which consists of 26 million product offers originating from 79 thousand websites.","description_withheld":null,"homepage":"http://webdatacommons.org/largescaleproductcorpus/v2/","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[{"name":"Tabular","url":"/datasets/modality/tabular"}],"tasks":[{"name":"Entity Resolution","url":"/task/entity-resolution","datasets_with_task":"/datasets/task/entity-resolution"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["WDC Watches-xlarge","WDC Watches-small","WDC Watches-medium","WDC Watches-large","WDC Shoes-xlarge","WDC Shoes-small","WDC Shoes-medium","WDC Shoes-large","WDC Cameras-xlarge","WDC Cameras-small","WDC Cameras-medium","WDC Cameras-large","WDC Computers-xlarge","WDC Computers-small","WDC Computers-medium","WDC Computers-large","WDC LSPM"],"data_loaders":[{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/wdc/products-2017","frameworks":["tf","pytorch","jax"]}],"num_papers_in_archive":8,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/entity-resolution-on-wdc-computers-small","task":"Entity Resolution","dataset_variant":"WDC Computers-small","rows":6,"metrics":["F1 (%)"],"first_row_in_archive_order":{"model":"BERT","paper":"/paper/intermediate-training-of-bert-for-product","metrics":{"F1 (%)":"96.53"},"code_links":[{"title":"weyoun2211/productbert-intermediate","url":"https://github.com/weyoun2211/productbert-intermediate"},{"title":"wbsg-uni-mannheim/productbert-intermediate","url":"https://github.com/wbsg-uni-mannheim/productbert-intermediate"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/entity-resolution-on-wdc-computers-xlarge","task":"Entity Resolution","dataset_variant":"WDC Computers-xlarge","rows":6,"metrics":["F1 (%)"],"first_row_in_archive_order":{"model":"RoBERTa-SupCon","paper":"/paper/supervised-contrastive-learning-for-product","metrics":{"F1 (%)":"98.33"},"code_links":[{"title":"wbsg-uni-mannheim/contrastive-product-matching","url":"https://github.com/wbsg-uni-mannheim/contrastive-product-matching"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/entity-resolution-on-wdc-watches-small","task":"Entity Resolution","dataset_variant":"WDC Watches-small","rows":4,"metrics":["F1 (%)"],"first_row_in_archive_order":{"model":"HG","paper":"/paper/entity-resolution-with-hierarchical-graph","metrics":{"F1 (%)":"94"},"code_links":[{"title":"CGCL-codes/HierGAT","url":"https://github.com/CGCL-codes/HierGAT"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/entity-resolution-on-wdc-watches-xlarge","task":"Entity Resolution","dataset_variant":"WDC Watches-xlarge","rows":3,"metrics":["F1 (%)"],"first_row_in_archive_order":{"model":"JointBERT","paper":"/paper/dual-objective-fine-tuning-of-bert-for-entity","metrics":{"F1 (%)":"97.09"},"code_links":[{"title":"wbsg-uni-mannheim/jointbert","url":"https://github.com/wbsg-uni-mannheim/jointbert"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/entity-resolution-with-hierarchical-graph","title":"Entity Resolution with Hierarchical Graph Attention Networks","date":"2022-06-01","rows_on_this_dataset":4,"code_links":1,"syntology":null},{"paper":"/paper/domain-adaptation-for-deep-entity-resolution","title":"Domain Adaptation for Deep Entity Resolution: A Design Space Exploration","date":"2022-06-01","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/supervised-contrastive-learning-for-product","title":"Supervised Contrastive Learning for Product Matching","date":"2022-02-04","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/dual-objective-fine-tuning-of-bert-for-entity","title":"Dual-Objective Fine-Tuning of BERT for Entity Matching","date":"2021-06-01","rows_on_this_dataset":4,"code_links":1,"syntology":null},{"paper":"/paper/profiling-entity-matching-benchmark-tasks","title":"Profiling Entity Matching Benchmark Tasks","date":"2020-10-19","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/intermediate-training-of-bert-for-product","title":"Intermediate Training of BERT for Product Matching","date":"2020-08-31","rows_on_this_dataset":2,"code_links":2,"syntology":null},{"paper":"/paper/deep-entity-matching-with-pre-trained","title":"Deep Entity Matching with Pre-Trained Language Models","date":"2020-04-01","rows_on_this_dataset":4,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":1,"samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}