Papers › Introducing a Family of Synthetic Datasets for Research on Bias in Machine Learning

Introducing a Family of Synthetic Datasets for Research on Bias in Machine Learning

19 Jul 2021arXiv:2107.08928archive 2025-07-28

William Blanzeisky, Pádraig Cunningham, Kenneth Kennedy

A significant impediment to progress in research on bias in machine learning (ML) is the availability of relevant datasets. This situation is unlikely to change much given the sensitivity of such data. For this reason, there is a role for synthetic data in this research. In this short paper, we present one such family of synthetic data sets. We provide an overview of the data, describe how the level of bias can be varied, and present a simple example of an experiment on the data.

PaperPDFCode

Code

williamblanzeisky/SBDG officialmentioned in papermentioned on GitHub report
williamblanzeisky/SyntheticDataGeneration officialmentioned in papermentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

BIG-bench Machine LearningSensitivity

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections