{"url":"/dataset/sap","name":"SAP","full_name":null,"description_markdown":"The **SAP benchmark** is a significant development in the realm of **attack prompt generation** for **red teaming** and **defending large language models (LLMs)**. Let's delve into the details:\r\n\r\n1. **Objective**:\r\n   - The primary goal of the SAP benchmark is to evaluate the safety and robustness of LLMs against **red teaming attacks**.\r\n   - Red teaming attacks involve inducing LLMs to generate harmful or inappropriate content.\r\n\r\n2. **Methodology**:\r\n   - The SAP benchmark combines both manual and automatic methods to generate high-quality attack prompts.\r\n   - It leverages the impressive capabilities of newly emerged LLMs.\r\n   - Specifically, it instructs LLMs to mimic human-generated prompts through **in-context learning**.\r\n   - The attack framework is designed to create these prompts.\r\n\r\n3. **Defense Framework**:\r\n   - In addition to attacking LLMs, the SAP benchmark proposes a defense framework.\r\n   - This framework fine-tunes victim LLMs through **iterative interactions** with the attack framework.\r\n   - The goal is to enhance the safety of LLMs against red teaming attacks.\r\n\r\n4. **Validation and Datasets**:\r\n   - Extensive experiments on different LLMs validate the effectiveness of both the attack and defense frameworks.\r\n   - As part of this work, the authors release a series of **attack prompt datasets** named **SAP** with varying sizes.\r\n   - These datasets facilitate safety evaluation and enhancement for a broader range of LLMs¹.","description_withheld":null,"homepage":"https://github.com/Aatrox103/SAP","introduced_date":"2023-10-19","introduced_date_note":null,"introduced_by":{"paper":"/paper/attack-prompt-generation-for-red-teaming-and","title":"Attack Prompt Generation for Red Teaming and Defending Large Language Models","first_author":"Boyi Deng","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["SAP"],"data_loaders":[],"num_papers_in_archive":13,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}