Papers › Fewer Errors, but More Stereotypes? The Effect of Model Size on Gender Bias

Fewer Errors, but More Stereotypes? The Effect of Model Size on Gender Bias

20 Jun 2022NAACL (GeBNLP) 2022 7arXiv:2206.09860archive 2025-07-28

Yarden Tal, Inbal Magar, Roy Schwartz

The size of pretrained models is increasing, and so is their performance on a variety of NLP tasks. However, as their memorization capacity grows, they might pick up more social biases. In this work, we examine the connection between model size and its gender bias (specifically, occupational gender bias). We measure bias in three masked language model families (RoBERTa, DeBERTa, and T5) in two setups: directly using prompt based method, and using a downstream task (Winogender). We find on the one hand that larger models receive higher bias scores on the former task, but when evaluated on the latter, they make fewer gender errors. To examine these potentially conflicting results, we carefully investigate the behavior of the different models on Winogender. We find that while larger models outperform smaller ones, the probability that their mistakes are caused by gender bias is higher. Moreover, we find that the proportion of stereotypical errors compared to anti-stereotypical ones grows with the model size. Our findings highlight the potential risks that can arise from increasing model size.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

schwartz-lab-nlp/model_size_and_gender_bias officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Language ModelingLanguage ModellingMemorization

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

DeBERTa

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections