Papers › StereoSet: Measuring stereotypical bias in pretrained language models
StereoSet: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, Siva Reddy
A stereotype is an over-generalized belief about a particular group of people, e.g., Asians are good at math or Asians are bad drivers. Such beliefs (biases) are known to hurt target groups. Since pretrained language models are trained on large real world data, they are known to capture stereotypical biases. In order to assess the adverse effects of these models, it is important to quantify the bias captured in them. Existing literature on quantifying bias evaluates pretrained language models on a small set of artificially constructed bias-assessing sentences. We present StereoSet, a large-scale natural dataset in English to measure stereotypical biases in four domains: gender, profession, race, and religion. We evaluate popular models like BERT, GPT-2, RoBERTa, and XLNet on our dataset and show that these models exhibit strong stereotypical biases. We also present a leaderboard with a hidden test set to track the bias of future language models at https://stereoset.mit.edu
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Datasets
Introduced by this paper, per the archive.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Bias Detection | StereoSet | GPT-2 (small) | ICAT Score | 72.97 | #1 of 11 | Archive leaderboard | report |
| Bias Detection | StereoSet | XLNet (large) | ICAT Score | 72.03 | #2 of 11 | Archive leaderboard | report |
| Bias Detection | StereoSet | GPT-2 (medium) | ICAT Score | 71.73 | #3 of 11 | Archive leaderboard | report |
| Bias Detection | StereoSet | BERT (base) | ICAT Score | 71.21 | #4 of 11 | Archive leaderboard | report |
| Bias Detection | StereoSet | GPT-2 (large) | ICAT Score | 70.54 | #5 of 11 | Archive leaderboard | report |
| Bias Detection | StereoSet | BERT (large) | ICAT Score | 69.89 | #6 of 11 | Archive leaderboard | report |
| Bias Detection | StereoSet | RoBERTa (base) | ICAT Score | 67.50 | #7 of 11 | Archive leaderboard | report |
| Bias Detection | StereoSet | XLNet (base) | ICAT Score | 62.10 | #9 of 11 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections