Datasets › WDC LSPM

WDC LSPM

archive 2025-07-28

Many e-shops have started to mark-up product data within their HTML pages using the schema.org vocabulary. The Web Data Commons project regularly extracts such data from the Common Crawl, a large public web crawl. The Web Data Commons Training and Test Sets for Large-Scale Product Matching contain product offers from different e-shops in the form of binary product pairs (with corresponding label "match" or "no match") for four product categories, computers, cameras, watches and shoes.

In order to support the evaluation of machine learning-based matching methods, the data is split into training, validation and test sets. For each product category, we provide training sets in four different sizes (2.000-70.000 pairs). Furthermore there are sets of ids for each training set for a possible validation split (stratified random draw) available. The test set for each product category consists of 1.100 product pairs. The labels of the test sets were manually checked while those of the training sets were derived using shared product identifiers from the Web via weak supervision.

The data stems from the WDC Product Data Corpus for Large-Scale Product Matching - Version 2.0 which consists of 26 million product offers originating from 79 thousand websites.

Benchmarks archive 2025-07-28

All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

7 shown of 7 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 8. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Entity Resolution with Hierarchical Graph Attention Networks 1 4 1 Jun 2022 not harvested
Domain Adaptation for Deep Entity Resolution: A Design Space Exploration 1 2 1 Jun 2022 not harvested
Supervised Contrastive Learning for Product Matching 1 2 4 Feb 2022 not harvested
Dual-Objective Fine-Tuning of BERT for Entity Matching 1 4 1 Jun 2021 not harvested
Profiling Entity Matching Benchmark Tasks 1 1 19 Oct 2020 not harvested
Intermediate Training of BERT for Product Matching 2 2 31 Aug 2020 not harvested
Deep Entity Matching with Pre-Trained Language Models 1 4 1 Apr 2020 ran 1 of 1 samples (0 unverified)

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • WDC Watches-xlarge
  • WDC Watches-small
  • WDC Watches-medium
  • WDC Watches-large
  • WDC Shoes-xlarge
  • WDC Shoes-small
  • WDC Shoes-medium
  • WDC Shoes-large
  • WDC Cameras-xlarge
  • WDC Cameras-small
  • WDC Cameras-medium
  • WDC Cameras-large
  • WDC Computers-xlarge
  • WDC Computers-small
  • WDC Computers-medium
  • WDC Computers-large
  • WDC LSPM

17 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections