Home › Datasets › task › Fairness

Fairness datasets

archive 2025-07-28

23 datasets carry the task tag "Fairness" (the task itself: Fairness), ordered by the archive's paper count. Page 1 of 1: 23 shown of 23. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Fairness datasets 1–23 of 23

Netflix Prize consists of about 100,000,000 ratings for 17,770 movies given by 480,189 users.
370 papers · 1 benchmark
The UTKFace dataset is a large-scale face dataset with long age span (range from 0 to 116 years old).
243 papers · 4 benchmarks
FairFace is a face image dataset which is race balanced.
208 papers · 1 benchmark
MORPH is a facial age estimation dataset, which contains 55,134 facial images of 13,617 subjects ranging from 16 to 77 years old.
180 papers · 8 benchmarks
WinoBias contains 3,160 sentences, split equally for development and test, created by researchers familiar with the project.
134 papers · 0 benchmarks
RFW (Racial Faces in-the-Wild)
To validate the racial bias of four commercial APIs and four state-of-the-art (SOTA) algorithms.
73 papers · 0 benchmarks
GVGAI (General Video Game AI)
The General Video Game AI (GVGAI) framework is widely used in research which features a corpus of over 100 single-player games and 60 two-player games.
36 papers · 0 benchmarks
The HELP dataset is an automatically created natural language inference (NLI) dataset that embodies the combination of lexical and logical inferences focusing on monotonicity (i.e., phrase replacement-based reasoning).
30 papers · 1 benchmark
MIAP (More Inclusive Annotations for People)
MIAP is a dataset created by obtaining a new set of annotations on a subset of the Open Images dataset, containing bounding boxes and attributes for all of the people visible in those images, as the original Open Images dataset annotations…
15 papers · 0 benchmarks
ACS PUMS stands for American Community Survey (ACS) Public Use Microdata Sample (PUMS) and has been used to construct several tabular datasets for studying fairness in machine learning: - ACSIncome: to predict whether an individual’s…
11 papers · 0 benchmarks
A new face annotation dataset with balanced distribution between genders and ethnic origins.
11 papers · 2 benchmarks
BAF (Bank Account Fraud)
Bank Account Fraud (BAF) is a large-scale, realistic suite of tabular datasets.
10 papers · 12 benchmarks
SPEECH-COCO contains speech captions that are generated using text-to-speech (TTS) synthesis resulting in 616,767 spoken captions (more than 600h) paired with images.
9 papers · 0 benchmarks
CI-MNIST (Correlated and Imbalanced MNIST)
CI-MNIST (Correlated and Imbalanced MNIST) is a variant of MNIST dataset with introduced different types of correlations between attributes, dataset features, and an artificial eligibility criterion.
4 papers · 0 benchmarks
This research aimed at the case of customers default payments in Taiwan and compares the predictive accuracy of probability of default among six data mining methods.
3 papers · 0 benchmarks
The Dialogue Fairness dataset is used to evaluate and understand fairness in dialogue models, focusing on gender and racial biases.
2 papers · 0 benchmarks
ImDrug is a comprehensive benchmark with an open-source Python library which consists of 4 imbalance settings, 11 AI-ready datasets, 54 learning tasks and 16 baseline algorithms tailored for imbalanced learning.
2 papers · 0 benchmarks
KANFace (KANFace Dataset)
KANFace consists of 40K still images and 44K sequences (14.5M video frames in total) captured in unconstrained, real-world conditions from 1,045 subjects.
2 papers · 1 benchmark
TwinViews-13k is a dataset of 13,855 pairs of left-leaning and right-leaning political statements, each pair matched by topic.
2 papers · 0 benchmarks
A Racial Fairness Benchmark Dataset for Face Forgery Detection.
1 paper · 0 benchmarks
ICLR Database (ICLR Database (with Textual Covariates))
A maintained database tracks ICLR submissions and reviews, augmented with author profiles and higher-level textual features.
1 paper · 0 benchmarks
ec-darkpattern is a dataset for dark pattern detection and prepared its baseline detection performance with state-of-the-art machine learning methods.
1 paper · 0 benchmarks
fNIRS2MW (The Tufts fNIRS to Mental Workload Dataset)
The Tufts fNIRS to Mental Workload (fNIRS2MW) open-access dataset is a new dataset for building machine learning classifiers that can consume a short window (30 seconds) of multivariate fNIRS recordings and predict the mental workload…
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.