Browse State-of-the-Art › Bias Detection
Bias Detection
80 papers with code · 5 benchmarks · 11 datasets archive 2025-07-28
Bias detection is the task of detecting and measuring racism, sexism and otherwise discriminatory behavior in a model (Source: https://stereoset.mit.edu/)
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
5 leaderboard tables shown for this task, 5 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| StereoSet (11 rows) | GPT-2 (small) | StereoSet: Measuring stereotypical bias in pretrained language models | code | — | Compare |
| rt-inod-bias (5 rows) | GPT-4 | Benchmarking Llama2, Mistral, Gemma and GPT for Factuality,... | code | Syntology ran 7 of 8 samples · 1 unverified | Compare |
| ICAT LLM bias (1 row) | BAD | BAD: BiAs Detection for Large Language Models in the context of... | code | — | Compare |
| PlantVillage_8px (1 row) | RandomForest_default_hyperparameters | Uncovering bias in the PlantVillage dataset | code | — | Compare |
| Wiki Neutrality Corpus (1 row) | RoBERTa+ALBERT | Towards Detection of Subjective Bias using Contextualized Word Embeddings | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
11 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 80 papers with code (199 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
20 Apr 2020 3 repositories listedSince pretrained language models are trained on large real world data, they are known to capture stereotypical biases.
-
16 Sep 2023 2 repositories listedLLMs are increasingly powerful and widely used to assist users in a variety of tasks.
-
28 Jan 2023 2 repositories listedHence, we also contribute a new, large Swedish bias-labelled dataset (of 2 million samples), translated from the English version and train the SotA mT5 model on it.
-
6 Jun 2020 2 repositories listed Syntology ran 6 of 10 samples · 4 unverified · 6 pointer-only (licence)Furthermore, we develop two methods, Intersectional Bias Detection (IBD) and Emergent Intersectional Bias Detection (EIBD), to automatically identify the intersectional biases and emergent intersectional biases from…
-
2 Dec 2019 2 repositories listedTo address these drawbacks, we formalize a method for automating the selection of interesting PDPs and extend PDPs beyond showing single features to show the model response along arbitrary directions, for example in raw…
-
1 Jun 2025 1 repository listedWe present Concept Trajectory Analysis (CTA), an interpretability method that tracks how neural networks organize concepts by following their paths through clustered activation spaces across layers.
-
19 May 2025 1 repository listedMedia bias detection is a critical task in ensuring fair and balanced information dissemination, yet it remains challenging due to the subjectivity of bias and the scarcity of high-quality annotated data.
-
16 May 2025 1 repository listedGenerative AI systems can help spread information but also misinformation and biases, potentially undermining the UN Sustainable Development Goals (SDGs).
-
1 May 2025 1 repository listedA key objective in spatial analysis is to model spatial relationships and infer spatial processes to generate knowledge from spatial data, which has been largely based on spatial statistical methods.
-
21 Feb 2025 1 repository listedThere is general agreement for the remaining 3 personality dimensions: both sides observe at most small differences across gender.
-
22 Dec 2024 1 repository listedThe integration of Large Language Models (LLMs) and Vision-Language Models (VLMs) opens new avenues for addressing complex challenges in multimodal content analysis, particularly in biased news detection.
-
16 Dec 2024 1 repository listedWe introduce MT-LENS, a framework designed to evaluate Machine Translation (MT) systems across a variety of tasks, including translation quality, gender bias detection, added toxicity, and robustness to misspellings.
-
17 Nov 2024 1 repository listedOur classifier, fine-tuned on this dataset, surpasses all of the annotator LLMs by 5-9 percent in Matthews Correlation Coefficient (MCC) and performs close to or outperforms the model trained on human-labeled data when…
-
Mitigating Bias in Queer Representation within Large Language Models: A Collaborative Agent Approach12 Nov 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedLarge Language Models (LLMs) often perpetuate biases in pronoun usage, leading to misrepresentation or exclusion of queer individuals.
-
17 Oct 2024 1 repository listedThis study displays the problems with the current benchmarks for measuring demographic bias in Vision Language Models and introduces both a more effective dataset for measuring bias and a novel and interpretable…
-
10 Oct 2024 1 repository listedThe detection of bias in natural language processing (NLP) is a critical challenge, particularly with the increasing use of large language models (LLMs) in various domains.
-
9 Oct 2024 1 repository listedOur approach features: (1) a synthetic emotional instruct dataset for both pre-training and fine-tuning stages, (2) a Metric Projector that delegates classification from the language model allowing for more efficient…
-
3 Oct 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)To that end, our study investigates the presence of implicit gender biases in multi-agent LLM interactions and proposes two strategies to mitigate these biases.
-
25 Sep 2024 1 repository listed Syntology ran 10 of 10 samples · 0 unverified · 10 pointer-only (licence)Our model allows any large language model to perform counterfactual token generation at almost no cost in comparison with vanilla token generation, it is embarrassingly simple to implement, and it does not require any…
-
4 Sep 2024 1 repository listedIn IDH mutation classification, HIPPO more robustly identified the pathology regions responsible for false negatives compared to attention, suggesting its potential to outperform attention in explaining model decisions.
-
29 Aug 2024 1 repository listedOpenBias detects and quantifies biases, while GradBias determines the contribution of individual prompt words on biases.
-
26 Jul 2024 1 repository listedThe project BIAS: Mitigating Diversity Biases of AI in the Labor Market is a four-year project funded by the European commission and supported by the Swiss State Secretariat for Education, Research and Innovation (SERI).
-
3 Jul 2024 1 repository listedThe rapid growth of Large Language Models (LLMs) has put forward the study of biases as a crucial field.
-
1 Jul 2024 1 repository listedIn this paper, we apply a method to quantify biases associated with named entities from various countries.
-
1 Jun 2024 1 repository listedData derived from the realm of the social sciences is often produced in digital text form, which motivates its use as a source for natural language processing methods.
-
20 May 2024 1 repository listedMachine Learning algorithms (ML) impact virtually every aspect of human lives and have found use across diverse sectors including healthcare, finance, and education.
-
16 May 2024 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedIt is often desirable to distill the capabilities of large language models (LLMs) into smaller student models due to compute and memory constraints.
-
15 Apr 2024 1 repository listed Syntology ran 7 of 8 samples · 1 unverifiedIn this research, we used OpenAI GPT as point of comparison since it excels at all levels of safety.
-
11 Apr 2024 1 repository listed Syntology ran 9 of 10 samples · 1 unverified · 10 pointer-only (licence)In this paper, we tackle the challenge of open-set bias detection in text-to-image generative models presenting OpenBias, a new pipeline that identifies and quantifies the severity of biases agnostically, without access…
-
28 Mar 2024 1 repository listed Syntology ran 6 of 7 samples · 1 unverifiedNeural networks trained on biased datasets tend to inadvertently learn spurious correlations, hindering generalization.
Syntology lines on 8 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections