{"url":"/dataset/redteaming-resistance-benchmark","name":"Redteaming Resistance Benchmark","full_name":null,"description_markdown":"The **Redteaming Resistance Benchmark** is a project aimed at evaluating the robustness of language models, both open-source and black-box, through **redteaming attacks**. These attacks involve systematically challenging and testing models with carefully crafted prompts to uncover their failure modes and vulnerabilities. In other words, it reveals where these models are susceptible to generating problematic outputs¹².\r\n\r\nHere are some key points about the **Redteaming Resistance Benchmark**:\r\n\r\n- **Purpose**: The project aims to identify weaknesses in language models by simulating adversarial scenarios.\r\n- **Datasets Used**: The benchmark uses several datasets to evaluate model performance, including:\r\n    - **Advbench**: A dataset of adversarial behaviors ranging from profanity, discrimination, to violence.\r\n    - **AART**: A collection of generated adversarial behaviors with diverse cultural, geographic, and application settings.\r\n    - **Beavertails**: A dataset for safety alignment research in large language models.\r\n    - **Do Not Answer**: A dataset of prompts to which responsible language models do not answer.\r\n    - **RedEval - HarmfulQA** and **RedEval - DangerousQA**: Datasets of harmful questions covering various topics.\r\n    - **Student-Teacher Prompting**: A dataset of harmful prompts that successfully broke a specific language model.\r\n    - **SAP**: Attacks generated through in-context learning to mimic human speech.\r\n- **Content Categories**: The benchmark evaluates model performance across 15 different categories, including illegal activity, harm, hate, discrimination, and more¹.\r\n\r\n(1) GitHub - haizelabs/redteaming-resistance-benchmark. https://github.com/haizelabs/redteaming-resistance-benchmark.\r\n(2) Introducing the Red-Teaming Resistance Leaderboard - Hugging Face. https://huggingface.co/blog/leaderboard-haizelab.\r\n(3) Redteaming Resistance Leaderboard - a Hugging Face Space by Nymbo. https://huggingface.co/spaces/Nymbo/red-teaming-resistance-benchmark.\r\n(4) Redteaming Resistance Benchmark - GitHub. https://github.com/jussker/redteaming-resistance-benchmark/blob/main/README.md.","description_withheld":null,"homepage":"https://github.com/haizelabs/redteaming-resistance-benchmark","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["Redteaming Resistance Benchmark"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}