{"url":"/dataset/gptfuzzer","name":"GPTFuzzer","full_name":null,"description_markdown":"**GPTFuzzer** is a fascinating project that explores **red teaming** of large language models (LLMs) using **auto-generated jailbreak prompts**. Let's dive into the details:\r\n\r\n1. **Project Overview**:\r\n   - **GPTFuzzer** aims to assess the security and robustness of LLMs by crafting prompts that can potentially lead to harmful or unintended behavior.\r\n   - The project focuses on **GPT-3** and similar models.\r\n\r\n2. **Datasets**:\r\n   - The datasets used in **GPTFuzzer** include:\r\n     - **Harmful Questions**: Sampled from public datasets like **llm-jailbreak-study** and **hh-rlhf**.\r\n     - **Human-Written Templates**: Collected from **llm-jailbreak-study**.\r\n     - **Responses**: Gathered by querying models like **Vicuna-7B**, **ChatGPT**, and **Llama-2-7B-chat**.\r\n\r\n3. **Models**:\r\n   - The judgment model is a **finetuned RoBERTa-large** model.\r\n   - The training code and data are available in the repository.\r\n   - During fuzzing experiments, the model is automatically downloaded and cached.\r\n\r\n4. **Updates**:\r\n   - The project has received recognition and awards at conferences like **Geekcon 2023**.\r\n   - The team continues to improve the codebase and aims to build a general black-box fuzzing framework for LLMs.\r\n\r\nSource: Conversation with Bing, 3/17/2024\r\n(1) sherdencooper/GPTFuzz: Official repo for GPTFUZZER - GitHub. https://github.com/sherdencooper/GPTFuzz.\r\n(2) GPTFUZZER : Red Teaming Large Language Models with Auto ... - GitHub. https://github.com/sherdencooper/GPTFuzz/blob/master/README.md.\r\n(3) GPTFUZZER : Red Teaming Large Language Models with Auto ... - GitHub. https://github.com/CriticalPulsar/GPTFuzz/blob/master/README.md.\r\n(4) undefined. https://avatars.githubusercontent.com/u/37368657?v=4.\r\n(5) undefined. https://github.com/sherdencooper/GPTFuzz/blob/master/README.md?raw=true.\r\n(6) undefined. https://desktop.github.com.\r\n(7) undefined. https://github.com/sherdencooper/GPTFuzz/raw/master/README.md.\r\n(8) undefined. https://opensource.org/licenses/MIT.\r\n(9) undefined. https://camo.githubusercontent.com/a4426cbe5c21edb002526331c7a8fbfa089e84a550567b02a0d829a98b136ad0/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4c6963656e73652d4d49542d79656c6c6f772e737667.\r\n(10) undefined. https://img.shields.io/badge/License-MIT-yellow.svg.\r\n(11) undefined. https://arxiv.org/pdf/2309.10253.pdf.\r\n(12) undefined. https://sherdencooper.github.io/.\r\n(13) undefined. https://scholar.google.com/citations?user=Zv_rC0AAAAAJ&amp.\r\n(14) undefined. http://www.dataisland.org/.\r\n(15) undefined. http://xinyuxing.org/.\r\n(16) undefined. https://geekcon.darknavy.com/2023/china/en/index.html.\r\n(17) undefined. https://avatars.githubusercontent.com/u/35443979?v=4.\r\n(18) undefined. https://github.com/CriticalPulsar/GPTFuzz/blob/master/README.md?raw=true.\r\n(19) undefined. https://docs.github.com/articles/about-issue-and-pull-request-templates.\r\n(20) undefined. https://github.com/CriticalPulsar/GPTFuzz/raw/master/README.md.\r\n(21) undefined. https://scholar.google.com/citations?user=Zv_rC0AAAAAJ&hl=en.","description_withheld":null,"homepage":"https://github.com/sherdencooper/GPTFuzz","introduced_date":"2023-09-19","introduced_date_note":null,"introduced_by":{"paper":"/paper/gptfuzzer-red-teaming-large-language-models","title":"GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts","first_author":"Jiahao Yu","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["GPTFuzzer"],"data_loaders":[],"num_papers_in_archive":44,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}