{"url":"/dataset/toxicchat","name":"ToxicChat","full_name":null,"description_markdown":"**ToxicChat** is a **novel benchmark dataset** constructed based on **real user queries** from an open-source chatbot. Unlike previous toxicity detection benchmarks that primarily rely on social media content, ToxicChat captures the **rich and nuanced phenomena** inherent in **real-world user-AI interactions**. This unique dataset reveals significant **domain differences** compared to social media contents, making it a valuable resource for exploring the challenges of toxicity detection in user-AI conversations¹.\r\n\r\nHere are some key details about the ToxicChat dataset:\r\n\r\n- **Construction**: ToxicChat was created using real user queries collected from an **open-source chatbot**.\r\n- **Challenges**: It contains phenomena that can be **tricky** for current toxicity detection models to identify.\r\n- **Domain Difference**: ToxicChat exhibits a significant **domain difference** when compared to social media content.\r\n- **Purpose**: ToxicChat serves as a benchmark to drive advancements in building a **safe and healthy environment** for user-AI interactions.\r\n\r\nSource: Conversation with Bing, 3/17/2024\r\n(1) ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real .... https://aclanthology.org/2023.findings-emnlp.311/.\r\n(2) arXiv:2310.17389v1 [cs.CL] 26 Oct 2023. https://arxiv.org/pdf/2310.17389.\r\n(3) README.md · lmsys/toxic-chat at main - Hugging Face. https://huggingface.co/datasets/lmsys/toxic-chat/blob/main/README.md.\r\n(4) The Toxicity Dataset - GitHub. https://github.com/surge-ai/toxicity.\r\n(5) undefined. https://aclanthology.org/2023.findings-emnlp.311.\r\n(6) undefined. https://aclanthology.org/2023.findings-emnlp.311.pdf.","description_withheld":null,"homepage":"https://huggingface.co/datasets/lmsys/toxic-chat","introduced_date":"2023-10-26","introduced_date_note":null,"introduced_by":{"paper":"/paper/toxicchat-unveiling-hidden-challenges-of","title":"ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation","first_author":"Zi Lin","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["ToxicChat"],"data_loaders":[],"num_papers_in_archive":34,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}