Browse State-of-the-Art › Red Teaming › Papers, page 2
Red Teaming
Papers archive 2025-07-28
archive papers tagged: 251 · with a code link: 110 · where Syntology ran a sample: 57 (45 with a run with no instrument failure, 12 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (57 of 251 tagged: 45 with a run with no instrument failure, 12 where every run was a failure of Syntology's instrument)
Page 2 of 3: papers 101 to 200 of 251, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
19 Oct 2023 1 repository listed Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
ASSERT: Automated Safety Scenario Red Teaming for Evaluating the Robustness of Large Language Models14 Oct 2023 1 repository listed
-
5 Oct 2023 1 repository listed Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
2 Oct 2023 1 repository listed
-
12 Sep 2023 1 repository listed Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
12 Aug 2023 1 repository listed Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
2 Aug 2023 1 repository listed Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
5 Jul 2023 1 repository listed
-
27 May 2023 1 repository listed Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
STACK: Adversarial Attacks on LLM Safeguard Pipelines30 Jun 2025 0 repositories listed
-
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models11 Jun 2025 0 repositories listed
-
Effective Red-Teaming of Policy-Adherent Agents11 Jun 2025 0 repositories listed
-
Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models8 Jun 2025 0 repositories listed
-
Red Teaming AI Policy: A Taxonomy of Avoision and the EU AI Act2 Jun 2025 0 repositories listed
-
A Red Teaming Roadmap Towards System-Level Safety30 May 2025 0 repositories listed
-
A Reward-driven Automated Webshell Malicious-code Generator for Red-teaming30 May 2025 0 repositories listed
-
Towards Secure MLOps: Surveying Attacks, Mitigation Strategies, and Research Challenges30 May 2025 0 repositories listed
-
CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring29 May 2025 0 repositories listed
-
SafeCOMM: What about Safety Alignment in Fine-Tuned Telecom Large Language Models?29 May 2025 0 repositories listed
-
Red-Teaming Text-to-Image Systems by Rule-based Preference Modeling27 May 2025 0 repositories listed
-
GhostPrompt: Jailbreaking Text-to-image Generative Models based on Dynamic Optimization25 May 2025 0 repositories listed
-
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation24 May 2025 0 repositories listed
-
Towards medical AI misalignment: a preliminary study22 May 2025 0 repositories listed
-
RRTL: Red Teaming Reasoning Large Language Models in Tool Learning21 May 2025 0 repositories listed
-
EVA: Red-Teaming GUI Agents via Evolving Indirect Prompt Injection20 May 2025 0 repositories listed
-
"Haet Bhasha aur Diskrimineshun": Phonetic Perturbations in Code-Mixed Hinglish to Red-Team LLMs20 May 2025 0 repositories listed
-
Hidden Ghost Hand: Unveiling Backdoor Vulnerabilities in MLLM-Powered Mobile GUI Agents20 May 2025 0 repositories listed
-
CURE: Concept Unlearning via Orthogonal Representation Editing in Diffusion Models19 May 2025 0 repositories listed
-
LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs16 May 2025 0 repositories listed
-
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents9 May 2025 0 repositories listed
-
Offensive Security for AI Systems: Concepts, Practices, and Applications9 May 2025 0 repositories listed
-
Safety by Measurement: A Systematic Literature Review of AI Safety Evaluation Methods8 May 2025 0 repositories listed
-
DMRL: Data- and Model-aware Reward Learning for Data Extraction7 May 2025 0 repositories listed
-
Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs7 May 2025 0 repositories listed
-
Red Teaming Large Language Models for Healthcare1 May 2025 0 repositories listed
-
When Testing AI Tests Us: Safeguarding Mental Health on the Digital Frontlines29 Apr 2025 0 repositories listed
-
RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models25 Apr 2025 0 repositories listed
-
Understanding and Mitigating Risks of Generative AI in Financial Services25 Apr 2025 0 repositories listed
-
ELAB: Extensive LLM Alignment Benchmark in Persian Language17 Apr 2025 0 repositories listed
-
X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents15 Apr 2025 0 repositories listed
-
Multi-lingual Multi-turn Automated Red Teaming for LLMs4 Apr 2025 0 repositories listed
-
Strategize Globally, Adapt Locally: A Multi-Turn Red Teaming Agent with Dual-Level Learning2 Apr 2025 0 repositories listed
-
Red Teaming with Artificial Intelligence-Driven Cyberattacks: A Scoping Review25 Mar 2025 0 repositories listed
-
AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration20 Mar 2025 0 repositories listed
-
MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models19 Mar 2025 0 repositories listed
-
A Framework for Evaluating Emerging Cyberattack Capabilities of AI14 Mar 2025 0 repositories listed
-
Making Every Step Effective: Jailbreaking Large Vision-Language Models Through Hierarchical KV Equalization14 Mar 2025 0 repositories listed
-
Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectives13 Mar 2025 0 repositories listed
-
JBFuzz: Jailbreaking LLMs Efficiently and Effectively Using Fuzzing12 Mar 2025 0 repositories listed
-
MAD-MAX: Modular And Diverse Malicious Attack MiXtures for Automated LLM Red Teaming8 Mar 2025 0 repositories listed
-
Reinforced Diffuser for Red Teaming Large Vision-Language Models8 Mar 2025 0 repositories listed
-
Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges6 Mar 2025 0 repositories listed
-
LLM-Safety Evaluations Lack Robustness4 Mar 2025 0 repositories listed
-
Building Safe GenAI Applications: An End-to-End Overview of Red Teaming for Large Language Models3 Mar 2025 0 repositories listed
-
Be a Multitude to Itself: A Prompt Evolution Framework for Red Teaming22 Feb 2025 0 repositories listed
-
Fast Proxies for LLM Robustness Evaluation14 Feb 2025 0 repositories listed
-
A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management10 Feb 2025 0 repositories listed
-
Predictive Red Teaming: Breaking Policies Without Breaking Robots10 Feb 2025 0 repositories listed
-
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs5 Feb 2025 0 repositories listed
-
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming31 Jan 2025 0 repositories listed
-
Playing Devil's Advocate: Unmasking Toxicity and Vulnerabilities in Large Vision-Language Models14 Jan 2025 0 repositories listed
-
Text-Diffusion Red-Teaming of Large Language Models: Unveiling Harmful Behaviors with Proximity Constraints14 Jan 2025 0 repositories listed
-
Lessons From Red Teaming 100 Generative AI Products13 Jan 2025 0 repositories listed
-
Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency9 Jan 2025 0 repositories listed
-
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models3 Jan 2025 0 repositories listed
-
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning24 Dec 2024 0 repositories listed
-
OpenAI o1 System Card21 Dec 2024 0 repositories listed
-
POEX: Understanding and Mitigating Policy Executable Jailbreak Attacks against Embodied AI21 Dec 2024 0 repositories listed
-
AI red-teaming is a sociotechnical challenge: on values, labor, and harms12 Dec 2024 0 repositories listed
-
Embodied Red Teaming for Auditing Robotic Foundation Models27 Nov 2024 0 repositories listed
-
In-Context Experience Replay Facilitates Safety Red-Teaming of Text-to-Image Diffusion Models25 Nov 2024 0 repositories listed
-
LLMStinger: Jailbreaking LLMs using RL fine-tuned LLMs13 Nov 2024 0 repositories listed
-
Desert Camels and Oil Sheikhs: Arab-Centric Red Teaming of Frontier LLMs31 Oct 2024 0 repositories listed
-
AdvAgent: Controllable Blackbox Red-teaming on Web Agents22 Oct 2024 0 repositories listed
-
LLM-Assisted Red Teaming of Diffusion Models through "Failures Are Fated, But Can Be Faded"22 Oct 2024 0 repositories listed
-
Insights and Current Gaps in Open-Source LLM Vulnerability Scanners: A Comparative Analysis21 Oct 2024 0 repositories listed
-
A Formal Framework for Assessing and Mitigating Emergent Security Risks in Generative AI Models: Bridging Theory and Dynamic Risk Mitigation15 Oct 2024 0 repositories listed
-
12 Oct 2024 0 repositories listed Syntology 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
Recent advancements in LLM Red-Teaming: Techniques, Defenses, and Ethical Considerations9 Oct 2024 0 repositories listed
-
SteerDiff: Steering towards Safe Text-to-Image Diffusion Models3 Oct 2024 0 repositories listed
-
Automated Red Teaming with GOAT: the Generative Offensive Agent Tester2 Oct 2024 0 repositories listed
-
Attack Atlas: A Practitioner's Perspective on Challenges and Pitfalls in Red Teaming GenAI23 Sep 2024 0 repositories listed
-
Jailbreaking Large Language Models with Symbolic Mathematics17 Sep 2024 0 repositories listed
-
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols12 Sep 2024 0 repositories listed
-
Exploring Straightforward Conversational Red-Teaming7 Sep 2024 0 repositories listed
-
Conversational Complexity for Assessing Risk in Large Language Models2 Sep 2024 0 repositories listed
-
Testing and Evaluation of Large Language Models: Correctness, Non-Toxicity, and Fairness31 Aug 2024 0 repositories listed
-
Atoxia: Red-teaming Large Language Models with Target Toxic Answers27 Aug 2024 0 repositories listed
-
DiffZOO: A Purely Query-Based Black-Box Attack for Red-teaming Text-to-Image Generative Model via Zeroth Order Optimization18 Aug 2024 0 repositories listed
-
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming14 Aug 2024 0 repositories listed
-
h4rm3l: A language for Composable Jailbreak Attack Synthesis9 Aug 2024 0 repositories listed
-
Can Large Language Models Automatically Jailbreak GPT-4V?23 Jul 2024 0 repositories listed
-
RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent23 Jul 2024 0 repositories listed
-
Breaking the Global North Stereotype: A Global South-centric Benchmark Dataset for Auditing and Mitigating Biases in Facial Recognition Systems22 Jul 2024 0 repositories listed
-
Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak Prompts21 Jul 2024 0 repositories listed
-
Phi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle18 Jul 2024 0 repositories listed
-
Direct Unlearning Optimization for Robust and Safe Text-to-Image Models17 Jul 2024 0 repositories listed
-
The Human Factor in AI Red Teaming: Perspectives from Social and Collaborative Computing10 Jul 2024 0 repositories listed
-
Purple-teaming LLMs with Adversarial Defender Training1 Jul 2024 0 repositories listed
Syntology lines on 8 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.