{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/rtp-lx-can-llms-evaluate-toxicity-in","title":"RTP-LX: Can LLMs Evaluate Toxicity in Multilingual Scenarios?","arxiv_id":"2404.14397","date":"2024-04-22","proceeding":null,"authors":["Adrian de Wynter","Ishaan Watts","Tua Wongsangaroonsri","Minghui Zhang","Noura Farra","Nektar Ege Altıntoprak","Lena Baur","Samantha Claudet","Pavel Gajdusek","Can Gören","Qilong Gu","Anna Kaminska","Tomasz Kaminski","Ruby Kuo","Akiko Kyuba","Jongho Lee","Kartik Mathur","Petter Merok","Ivana Milovanović","Nani Paananen","Vesa-Matti Paananen","Anna Pavlenko","Bruno Pereira Vidal","Luciano Strika","Yueh Tsao","Davide Turcato","Oleksandr Vakhno","Judit Velcsov","Anna Vickers","Stéphanie Visser","Herdyan Widarmanto","Andrey Zaikin","Si-Qing Chen"],"abstract":"Large language models (LLMs) and small language models (SLMs) are being adopted at remarkable speed, although their safety still remains a serious concern. With the advent of multilingual S/LLMs, the question now becomes a matter of scale: can we expand multilingual safety evaluations of these models with the same velocity at which they are deployed? To this end, we introduce RTP-LX, a human-transcreated and human-annotated corpus of toxic prompts and outputs in 28 languages. RTP-LX follows participatory design practices, and a portion of the corpus is especially designed to detect culturally-specific toxic language. We evaluate 10 S/LLMs on their ability to detect toxic content in a culturally-sensitive, multilingual scenario. We find that, although they typically score acceptably in terms of accuracy, they have low agreement with human judges when scoring holistically the toxicity of a prompt; and have difficulty discerning harm in context-dependent scenarios, particularly with subtle-yet-harmful content (e.g. microaggressions, bias). We release this dataset to contribute to further reduce harmful uses of these models and improve their safe deployment.","url_abs":"https://arxiv.org/abs/2404.14397v2","url_pdf":"https://arxiv.org/pdf/2404.14397v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"rtp-lx-can-llms-evaluate-toxicity-in","repo_url":"https://github.com/microsoft/rtp-lx","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2404.14397","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}