{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/one-language-many-gaps-evaluating-dialect","title":"One Language, Many Gaps: Evaluating Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks","arxiv_id":"2410.11005","date":"2024-10-14","proceeding":null,"authors":["Fangru Lin","Shaoguang Mao","Emanuele La Malfa","Valentin Hofmann","Adrian de Wynter","Xun Wang","Si-Qing Chen","Michael Wooldridge","Janet B. Pierrehumbert","Furu Wei"],"abstract":"Language is not monolithic. While benchmarks, including those designed for multiple languages, are often used as proxies to evaluate the performance of Large Language Models (LLMs), they tend to overlook the nuances of within-language variation, and thus fail to model the experience of speakers of non-standard dialects. Focusing on African American Vernacular English (AAVE), we present the first study aimed at objectively assessing the fairness and robustness of LLMs in handling dialects in canonical reasoning tasks, including algorithm, math, logic, and integrated reasoning. We introduce \\textbf{ReDial} (\\textbf{Re}asoning with \\textbf{Dial}ect Queries), a benchmark containing 1.2K+ parallel query pairs in Standardized English and AAVE. We hire AAVE speakers, including experts with computer science backgrounds, to rewrite seven popular benchmarks, such as HumanEval and GSM8K. With ReDial, we evaluate widely used LLMs, including GPT, Claude, Llama, Mistral, and the Phi model families. Our findings reveal that \\textbf{almost all of these widely used models show significant brittleness and unfairness to queries in AAVE}. Our work establishes a systematic and objective framework for analyzing LLM bias in dialectal queries. Moreover, it highlights how mainstream LLMs provide unfair service to dialect speakers in reasoning tasks, laying a critical foundation for relevant future research. Code and data can be accessed at https://github.com/fangru-lin/redial_dialect_robustness_fairness.","url_abs":"https://arxiv.org/abs/2410.11005v2","url_pdf":"https://arxiv.org/pdf/2410.11005v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"one-language-many-gaps-evaluating-dialect","repo_url":"https://github.com/fangru-lin/redial_dialect_robustness_fairness","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"fairness","task_name":"Fairness"},{"task_slug":"gsm8k","task_name":"GSM8K"},{"task_slug":"humaneval","task_name":"HumanEval"},{"task_slug":"math","task_name":"Math"}],"methods":[{"method_slug":"adam","method_name":"Adam"},{"method_slug":null,"method_name":"American"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"attention-dropout","method_name":"Attention Dropout"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"cosine-annealing","method_name":"Cosine Annealing"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"discriminative-fine-tuning","method_name":"Discriminative Fine-Tuning"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"gpt","method_name":"GPT"},{"method_slug":null,"method_name":null},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"linear-warmup-with-cosine-annealing","method_name":"Linear Warmup With Cosine Annealing"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"weight-decay","method_name":"Weight Decay"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2410.11005","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2410.11005"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/fangru-lin/redial_dialect_robustness_fairness","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":4,"unverified":3},"by_repo_kind":{"official":{"samples":7,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"59f2900e7b1f8bc2","entry":"find_answer_comprehensive","repo":"fangru-lin/redial_dialect_robustness_fairness","repo_kind":"official","path":"utils/utils.py","file_url":"https://github.com/fangru-lin/redial_dialect_robustness_fairness/blob/HEAD/utils/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"59f2900e7b1f8bc2"}},{"code_sha256_prefix":"ad59a66e901fd34f","entry":"generate_from_openai_chat_completion","repo":"fangru-lin/redial_dialect_robustness_fairness","repo_kind":"official","path":"utils/llm_call.py","file_url":"https://github.com/fangru-lin/redial_dialect_robustness_fairness/blob/HEAD/utils/llm_call.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"ad59a66e901fd34f"}},{"code_sha256_prefix":"6b96fd4ade51244c","entry":"inference","repo":"fangru-lin/redial_dialect_robustness_fairness","repo_kind":"official","path":"utils/open_llm_call.py","file_url":"https://github.com/fangru-lin/redial_dialect_robustness_fairness/blob/HEAD/utils/open_llm_call.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6b96fd4ade51244c"}},{"code_sha256_prefix":"7a893a21638d9a89","entry":"str_to_timedelta_list","repo":"fangru-lin/redial_dialect_robustness_fairness","repo_kind":"official","path":"utils/utils.py","file_url":"https://github.com/fangru-lin/redial_dialect_robustness_fairness/blob/HEAD/utils/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"7a893a21638d9a89"}},{"code_sha256_prefix":"702a98c04d4c24bb","entry":"check_correctness_comprehensive","repo":"fangru-lin/redial_dialect_robustness_fairness","repo_kind":"official","path":"utils/utils.py","file_url":"https://github.com/fangru-lin/redial_dialect_robustness_fairness/blob/HEAD/utils/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"702a98c04d4c24bb"}},{"code_sha256_prefix":"1b20637da1c0aa54","entry":"init_pipeline","repo":"fangru-lin/redial_dialect_robustness_fairness","repo_kind":"official","path":"utils/open_llm_call.py","file_url":"https://github.com/fangru-lin/redial_dialect_robustness_fairness/blob/HEAD/utils/open_llm_call.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"1b20637da1c0aa54"}},{"code_sha256_prefix":"f00384bec73e9e28","entry":"pipe_and_infer","repo":"fangru-lin/redial_dialect_robustness_fairness","repo_kind":"official","path":"utils/open_llm_call.py","file_url":"https://github.com/fangru-lin/redial_dialect_robustness_fairness/blob/HEAD/utils/open_llm_call.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f00384bec73e9e28"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}