{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/measuring-implicit-bias-in-explicitly","title":"Measuring Implicit Bias in Explicitly Unbiased Large Language Models","arxiv_id":"2402.04105","date":"2024-02-06","proceeding":null,"authors":["Xuechunzi Bai","Angelina Wang","Ilia Sucholutsky","Thomas L. Griffiths"],"abstract":"Large language models (LLMs) can pass explicit social bias tests but still harbor implicit biases, similar to humans who endorse egalitarian beliefs yet exhibit subtle biases. Measuring such implicit biases can be a challenge: as LLMs become increasingly proprietary, it may not be possible to access their embeddings and apply existing bias measures; furthermore, implicit biases are primarily a concern if they affect the actual decisions that these systems make. We address both challenges by introducing two new measures of bias: LLM Implicit Bias, a prompt-based method for revealing implicit bias; and LLM Decision Bias, a strategy to detect subtle discrimination in decision-making tasks. Both measures are based on psychological research: LLM Implicit Bias adapts the Implicit Association Test, widely used to study the automatic associations between concepts held in human minds; and LLM Decision Bias operationalizes psychological results indicating that relative evaluations between two candidates, not absolute evaluations assessing each independently, are more diagnostic of implicit biases. Using these measures, we found pervasive stereotype biases mirroring those in society in 8 value-aligned models across 4 social categories (race, gender, religion, health) in 21 stereotypes (such as race and criminality, race and weapons, gender and science, age and negativity). Our prompt-based LLM Implicit Bias measure correlates with existing language model embedding-based bias methods, but better predicts downstream behaviors measured by LLM Decision Bias. These new prompt-based measures draw from psychology's long history of research into measuring stereotype biases based on purely observable behavior; they expose nuanced biases in proprietary value-aligned LLMs that appear unbiased according to standard benchmarks.","url_abs":"https://arxiv.org/abs/2402.04105v2","url_pdf":"https://arxiv.org/pdf/2402.04105v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"measuring-implicit-bias-in-explicitly","repo_url":"https://github.com/baixuechunzi/llm-implicit-bias","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"measuring-implicit-bias-in-explicitly","repo_url":"https://github.com/ucabcg3/msc_bias_llm_project","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"diagnostic","task_name":"Diagnostic"},{"task_slug":"language-modelling","task_name":"Language Modelling"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2402.04105","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2402.04105"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/baixuechunzi/llm-implicit-bias","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ucabcg3/msc_bias_llm_project","reach":{"status":"ok"}}],"summary":{"ran_violates":2,"ran_honours":1},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"ef8f74084d195a43","entry":"cos_sim","repo":"baixuechunzi/llm-implicit-bias","repo_kind":"official","path":"analysis/weat.py","file_url":"https://github.com/baixuechunzi/llm-implicit-bias/blob/HEAD/analysis/weat.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"ef8f74084d195a43"}},{"code_sha256_prefix":"371636bf8dc99e05","entry":"unit_vector","repo":"baixuechunzi/llm-implicit-bias","repo_kind":"official","path":"analysis/weat.py","file_url":"https://github.com/baixuechunzi/llm-implicit-bias/blob/HEAD/analysis/weat.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"371636bf8dc99e05"}},{"code_sha256_prefix":"7adb0412e288ef26","entry":"weat_association","repo":"baixuechunzi/llm-implicit-bias","repo_kind":"official","path":"analysis/weat.py","file_url":"https://github.com/baixuechunzi/llm-implicit-bias/blob/HEAD/analysis/weat.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"7adb0412e288ef26"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}