Datasets › WDC Block

WDC Block (WDC Block: A Blocking Benchmark)

Introduced by Alexander Brinkmann et al. in SC-Block: Supervised Contrastive Blocking within Entity Resolution Pipelines6 Mar 2023 archive 2025-07-28

WDC Block is a benchmark for comparing the performance of blocking methods that are used as part of entity resolution pipelines.

Entity resolution aims to identify records in two datasets (A and B) that describe the same real-world entity. Since comparing all record pairs between two datasets can be computationally expensive, entity resolution is approached in two steps, blocking and matching. Blocking applies a computationally cheap method to remove non-matching record pairs and produces a smaller set of candidate record pairs reducing the workload of the matcher. During matching a more expensive pair-wise matcher produces a final set of matching record pairs.

Existing benchmark datasets for blocking and matching are rather small with respect to the Cartesian product AxB for comparing all records and the vocabulary size. If blockers are evaluated only on these small datasets, effects resulting from a high number of records or from a large vocabulary size (large number of unique tokens that need to be indexed) may be missed. The Web Data Commons Block (WDC-Block) is a new blocking benchmark that provides much larger datasets and thus requires blockers that address these scalability challenges. WDC Block features a maximal Cartesian product of 200 billion pairs of product offers which were extracted form 3,259 e-shops. Additionally, we provide three development sets with different sizes (~1K pairs, ~5K pairs & ~20K pairs) to experiment with different amounts of training data for the blockers.

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Blocking WDC Block - large SC-Block Candidate Set Size 5000000 SC-Block: Supervised Contrastive Blocking within Entity... wbsg-uni-mannheim/sc-block 2 Compare
Blocking WDC Block - medium SC-Block Candidate Set Size 100000 SC-Block: Supervised Contrastive Blocking within Entity... wbsg-uni-mannheim/sc-block 2 Compare
Blocking WDC Block - small BM25 Recall 96.9% SC-Block: Supervised Contrastive Blocking within Entity... wbsg-uni-mannheim/sc-block 2 Compare

Papers archive 2025-07-28

1 shown of 1 paper with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
SC-Block: Supervised Contrastive Blocking within Entity Resolution Pipelines 1 6 6 Mar 2023 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • WDC Block
  • WDC Block - small
  • WDC Block - medium
  • WDC Block - large

4 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections