Papers › StereoSet: Measuring stereotypical bias in pretrained language models

StereoSet: Measuring stereotypical bias in pretrained language models

20 Apr 2020ACL 2021 5arXiv:2004.09456archive 2025-07-28

Moin Nadeem, Anna Bethke, Siva Reddy

A stereotype is an over-generalized belief about a particular group of people, e.g., Asians are good at math or Asians are bad drivers. Such beliefs (biases) are known to hurt target groups. Since pretrained language models are trained on large real world data, they are known to capture stereotypical biases. In order to assess the adverse effects of these models, it is important to quantify the bias captured in them. Existing literature on quantifying bias evaluates pretrained language models on a small set of artificially constructed bias-assessing sentences. We present StereoSet, a large-scale natural dataset in English to measure stereotypical biases in four domains: gender, profession, race, and religion. We evaluate popular models like BERT, GPT-2, RoBERTa, and XLNet on our dataset and show that these models exhibit strong stereotypical biases. We also present a leaderboard with a hidden test set to track the bias of future language models at https://stereoset.mit.edu

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

moinnadeem/StereoSet officialpytorchCC-BY-SA-4.0 report
kanekomasahiro/evaluate_bias_in_mlm mentioned on GitHubpytorch report
zalkikar/mlm-bias mentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Bias DetectionMath

Datasets

Introduced by this paper, per the archive.

StereoSet

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Bias Detection StereoSet GPT-2 (small) ICAT Score 72.97 #1 of 11 Archive leaderboard report
Bias Detection StereoSet XLNet (large) ICAT Score 72.03 #2 of 11 Archive leaderboard report
Bias Detection StereoSet GPT-2 (medium) ICAT Score 71.73 #3 of 11 Archive leaderboard report
Bias Detection StereoSet BERT (base) ICAT Score 71.21 #4 of 11 Archive leaderboard report
Bias Detection StereoSet GPT-2 (large) ICAT Score 70.54 #5 of 11 Archive leaderboard report
Bias Detection StereoSet BERT (large) ICAT Score 69.89 #6 of 11 Archive leaderboard report
Bias Detection StereoSet RoBERTa (base) ICAT Score 67.50 #7 of 11 Archive leaderboard report
Bias Detection StereoSet XLNet (base) ICAT Score 62.10 #9 of 11 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AdamAttentionAttention DropoutBERTBPECosine AnnealingDense ConnectionsDiscriminative Fine-TuningDropoutGPT-2Layer NormalizationLinear LayerLinear Warmup With Cosine AnnealingLinear Warmup With Linear DecayMulti-Head AttentionResidual ConnectionRoBERTaSentencePieceSoftmaxWeight DecayWordPieceXLNet

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections