Datasets › WDC Block
WDC Block (WDC Block: A Blocking Benchmark)
WDC Block is a benchmark for comparing the performance of blocking methods that are used as part of entity resolution pipelines.
Entity resolution aims to identify records in two datasets (A and B) that describe the same real-world entity. Since comparing all record pairs between two datasets can be computationally expensive, entity resolution is approached in two steps, blocking and matching. Blocking applies a computationally cheap method to remove non-matching record pairs and produces a smaller set of candidate record pairs reducing the workload of the matcher. During matching a more expensive pair-wise matcher produces a final set of matching record pairs.
Existing benchmark datasets for blocking and matching are rather small with respect to the Cartesian product AxB for comparing all records and the vocabulary size. If blockers are evaluated only on these small datasets, effects resulting from a high number of records or from a large vocabulary size (large number of unique tokens that need to be indexed) may be missed. The Web Data Commons Block (WDC-Block) is a new blocking benchmark that provides much larger datasets and thus requires blockers that address these scalability challenges. WDC Block features a maximal Cartesian product of 200 billion pairs of product offers which were extracted form 3,259 e-shops. Additionally, we provide three development sets with different sizes (~1K pairs, ~5K pairs & ~20K pairs) to experiment with different amounts of training data for the blockers.
Benchmarks archive 2025-07-28
All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Blocking | WDC Block - large | SC-Block Candidate Set Size 5000000 | SC-Block: Supervised Contrastive Blocking within Entity... | wbsg-uni-mannheim/sc-block | 2 | Compare |
| Blocking | WDC Block - medium | SC-Block Candidate Set Size 100000 | SC-Block: Supervised Contrastive Blocking within Entity... | wbsg-uni-mannheim/sc-block | 2 | Compare |
| Blocking | WDC Block - small | BM25 Recall 96.9% | SC-Block: Supervised Contrastive Blocking within Entity... | wbsg-uni-mannheim/sc-block | 2 | Compare |
Papers archive 2025-07-28
1 shown of 1 paper with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| SC-Block: Supervised Contrastive Blocking within Entity Resolution Pipelines | 1 | 6 | 6 Mar 2023 | not harvested |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- WDC Block
- WDC Block - small
- WDC Block - medium
- WDC Block - large
4 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections