Browse State-of-the-Art › Red Teaming
Red Teaming
110 papers with code · 1 benchmark · 1 dataset archive 2025-07-28
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| SUDO Dataset (1 row) | SUDO | sudo rm -rf agentic_security | code | Syntology ran 2 of 2 samples · 0 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 110 papers with code (251 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
26 Jun 2024 3 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)As WildJailbreak considerably upgrades the quality and scale of existing safety resources, it uniquely enables us to examine the scaling effects of data and the interplay of data properties and model capabilities during…
-
6 Feb 2024 3 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedAutomated red teaming holds substantial promise for uncovering and mitigating the risks associated with the malicious use of large language models (LLMs), yet the field lacks a standardized evaluation framework to…
-
19 Sep 2023 3 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedRemarkably, GPTFuzz achieves over 90% attack success rates against ChatGPT and Llama-2 models, even with suboptimal initial seed templates.
-
15 Jun 2023 3 repositories listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)Using a pre-existing classifier does not allow for red-teaming to be tailored to the target model.
-
3 Oct 2024 2 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedIn this paper, we propose AutoDAN-Turbo, a black-box jailbreak method that can automatically discover as many jailbreak strategies as possible from scratch, without any human intervention or predefined scopes (e.
-
1 Aug 2024 2 repositories listed Syntology ran 9 of 16 samples · 7 unverifiedRapid advances in the capabilities of large language models (LLMs) have raised widespread concerns regarding their potential for malicious use.
-
22 Jul 2024 2 repositories listed Syntology ran 5 of 11 samples · 6 unverifiedFor example, the LLM red-teaming literature has produced a wide variety of 'jailbreaking' techniques to elicit harmful text from models that were fine-tuned to be harmless.
-
16 Jun 2024 2 repositories listedAs Large Language Models (LLMs) are deployed and integrated into thousands of applications, the need for scalable evaluation of how models respond to adversarial attacks grows rapidly.
-
6 Apr 2024 2 repositories listedWhen building Large Language Models (LLMs), it is paramount to bear safety in mind and protect them with guardrails.
-
8 Mar 2024 2 repositories listedIn this work, we utilize latent adversarial training (LAT) to defend against vulnerabilities without leveraging knowledge of what they are or using inputs that elicit them.
-
7 Mar 2024 2 repositories listed Syntology ran 4 of 4 samples · 0 unverifiedOur recipe for training the aligner models solely relies on synthetic data generated with a (prompted) LLM and can be easily adjusted for a variety of alignment criteria.
-
16 Oct 2023 2 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)While efforts have been made to mitigate such problems, either by implementing a safety filter at the evaluation stage or by fine-tuning models to eliminate undesirable concepts or styles, the effectiveness of these…
-
14 Oct 2023 2 repositories listedTo the best of our knowledge, our work is among the first to explore LLM unlearning.
-
10 Oct 2023 2 repositories listed Syntology ran 5 of 6 samples · 1 unverified · 6 pointer-only (licence)Finally, we propose an effective alignment method that explores diverse generation strategies, which can reasonably reduce the misalignment rate under our attack.
-
18 Aug 2023 2 repositories listedIn this work, we propose a new safety evaluation benchmark RED-EVAL that carries out red-teaming.
-
31 May 2023 2 repositories listed Syntology ran 1 of 18 samples · 17 unverifiedThe prevalence and strong capability of large language models (LLMs) present significant safety and ethical risks if exploited by malicious users.
-
5 Sep 2022 2 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)In this work, we study white-box adversarial policies and show that having access to a target agent's internal state can be useful for identifying its vulnerabilities.
-
23 Aug 2022 2 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedWe provide our own analysis of the data and find a variety of harmful outputs, which range from offensive language to more subtly harmful non-violent unethical outputs.
-
8 Jul 2025 1 repository listedLarge language models (LLMs) and their safety classifiers often perform poorly on low-resource languages due to limited training data and evaluation benchmarks.
-
16 Jun 2025 1 repository listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)In this position paper, we advocate the research community in LLM safety to pay close attention to the new safety risks issues introduced by MCP, and develop new techniques to build safe MCP-powered agent systems.
-
4 Jun 2025 1 repository listedRed teaming has proven to be an effective method for identifying and mitigating vulnerabilities in Large Language Models (LLMs).
-
4 Jun 2025 1 repository listedWe propose RedDebate, a novel multi-agent debate framework that leverages adversarial argumentation among Large Language Models (LLMs) to proactively identify and mitigate their own unsafe behaviours.
-
3 Jun 2025 1 repository listedThe inherent risk of generating harmful and unsafe content by Large Language Models (LLMs), has highlighted the need for their safety alignment.
-
30 May 2025 1 repository listedLarge Language Models (LLMs) excel in various natural language processing tasks but remain vulnerable to generating harmful content or being exploited for malicious purposes.
-
28 May 2025 1 repository listed Syntology ran 3 of 17 samples · 14 unverifiedNevertheless, we observe concerning ASRs of up to 50% in realistic end-to-end settings, with the recently released frontier Claude 4 Opus | CUA showing an alarming ASR of 48%, demonstrating that indirect prompt…
-
26 May 2025 1 repository listed Syntology ran 0 of 23 samples · 23 unverifiedThree strong trends emerge: (i) more capable models are better attackers, (ii) attack success drops sharply once the target's capability exceeds the attacker's, and (iii) attack success rates correlate with high…
-
22 May 2025 1 repository listedIn the adversarial iterative optimization stage, the red-team model and the target model continuously improve their respective capabilities in interaction.
-
20 May 2025 1 repository listedTo help evaluate and understand the latent capabilities of language models, this paper introduces an approach using optimized input embeddings, or 'soft prompts,' as a metric of conditional distance between a model and…
-
11 May 2025 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Our findings reveal that fine-tuning LLMs on 100 outlier samples selected by Self-Inf-N in the benign datasets severely compromises LLM safety alignment.
-
1 May 2025 1 repository listedLarge Language Models (LLMs) have demonstrated remarkable capabilities in natural language understanding and generation, enabling their widespread adoption across various domains.
Syntology lines on 17 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections