{"url":"/dataset/advsuffixes","name":"AdvSuffixes","full_name":"Adversarial Suffixes","description_markdown":"# AdvSuffixes - Information\r\n\r\nAdvSuffixes is a curated dataset of adversarial prompts and suffixes designed to evaluate and enhance the robustness of large language models (LLMs) against adversarial attacks. By appending these suffixes to standard prompts, researchers and developers can explore and analyze how LLMs respond to potentially harmful input scenarios. This dataset is heavily inspired by [AdvBench](https://github.com/llm-attacks/llm-attacks).\r\n\r\n## Dataset Structure\r\nThe dataset is organized as follows:\r\n```\r\ndata/\r\n│\r\n├── advsuffixes/\r\n│   ├── advsuffixes.csv      # Adversarial suffixes and their respective prompts\r\n│   ├── advsuffixes_eval.txt # 100 additional evaluation prompts that are out-of-distribution\r\n│\r\n├── ...\r\n```\r\n\r\n- `advsuffixes.csv`: The primary dataset containing pairs of adversarial suffixes and corresponding prompts.  \r\n- `advsuffixes_eval.txt`: A set of 100 additional evaluation prompts designed to test model robustness, that are out-of-distribution from the original 519 prompts in `advsuffixes.csv`. \r\n\r\nDetails about the dataset generation have been provided in Appendix B in the supplementary materials of the paper. There are 11763 listed suffixes overall, averaging 22.6 suffixes per prompt.\r\n\r\n---\r\n\r\n## License\r\nThis dataset is distributed under the GNU General Public License v3.0.","description_withheld":null,"homepage":"https://github.com/TrustMLRG/GASP/tree/main/data/advsuffixes","introduced_date":"2024-11-21","introduced_date_note":null,"introduced_by":{"paper":"/paper/gasp-efficient-black-box-generation-of","title":"GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMs","first_author":"Advik Raj Basani","url":null},"license":{"name":"GPL-3.0 License","url":"https://github.com/TrustMLRG/GASP/blob/main/LICENSE"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Adversarial Robustness","url":"/task/adversarial-robustness","datasets_with_task":"/datasets/task/adversarial-robustness"},{"name":"Adversarial Attack","url":"/task/adversarial-attack","datasets_with_task":"/datasets/task/adversarial-attack"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["AdvSuffixes"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}