Browse State-of-the-Art › Red Teaming › Papers, page 3
Red Teaming
Papers archive 2025-07-28
archive papers tagged: 251 · with a code link: 110 · where Syntology ran a sample: 57 (45 with a run with no instrument failure, 12 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (57 of 251 tagged: 45 with a run with no instrument failure, 12 where every run was a failure of Syntology's instrument)
Page 3 of 3: papers 201 to 251 of 251, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm26 Jun 2024 0 repositories listed
-
Leveraging Reinforcement Learning in Red Teaming for Advanced Ransomware Attack Simulations25 Jun 2024 0 repositories listed
-
Adversaries Can Misuse Combinations of Safe Models20 Jun 2024 0 repositories listed
-
20 Jun 2024 0 repositories listed Syntology 14 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 14 harvested samples) · 1 pointer-only (licence)
-
CELL your Model: Contrastive Explanations for Large Language Models17 Jun 2024 0 repositories listed
-
Ruby Teaming: Improving Quality Diversity Search with Memory for Automated Red Teaming17 Jun 2024 0 repositories listed
-
STAR: SocioTechnical Approach to Red Teaming Language Models17 Jun 2024 0 repositories listed
-
Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters30 May 2024 0 repositories listed
-
Safety Alignment for Vision Language Models22 May 2024 0 repositories listed
-
Tiny Refinements Elicit Resilience: Toward Efficient Prefix-Model Against LLM Red-Teaming21 May 2024 0 repositories listed
-
A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI23 Apr 2024 0 repositories listed
-
CulturalTeaming: AI-Assisted Interactive Red-Teaming for Challenging LLMs' (Lack of) Multicultural Knowledge10 Apr 2024 0 repositories listed
-
Aurora-M: Open Source Continual Pre-training for Multilingual Language and Code30 Mar 2024 0 repositories listed
-
IterAlign: Iterative Constitutional Alignment of Large Language Models27 Mar 2024 0 repositories listed
-
HRLAIF: Improvements in Helpfulness and Harmlessness in Open-domain Reinforcement Learning From AI Feedback13 Mar 2024 0 repositories listed
-
A Safe Harbor for AI Evaluation and Red Teaming7 Mar 2024 0 repositories listed
-
AttackGNN: Red-Teaming GNNs in Hardware Security Using Reinforcement Learning21 Feb 2024 0 repositories listed
-
Investigating Bias Representations in Llama 2 Chat via Activation Steering1 Feb 2024 0 repositories listed
-
Red-Teaming for Generative AI: Silver Bullet or Security Theater?29 Jan 2024 0 repositories listed
-
Towards Red Teaming in Multimodal and Multilingual Translation29 Jan 2024 0 repositories listed
-
Digital cloning of online social networks for language-sensitive agent-based modeling of misinformation spread23 Jan 2024 0 repositories listed
-
Red Teaming Visual Language Models23 Jan 2024 0 repositories listed
-
A Red Teaming Framework for Securing AI in Maritime Autonomous Systems8 Dec 2023 0 repositories listed
-
DeceptPrompt: Exploiting LLM-driven Code Generation via Adversarial Natural Language Instructions7 Dec 2023 0 repositories listed
-
InfoPattern: Unveiling Information Propagation Patterns in Social Media27 Nov 2023 0 repositories listed
-
JAB: Joint Adversarial Prompting and Belief Augmentation16 Nov 2023 0 repositories listed
-
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models16 Nov 2023 0 repositories listed
-
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts15 Nov 2023 0 repositories listed
-
Towards Publicly Accountable Frontier LLMs: Building an External Scrutiny Ecosystem under the ASPIRE Framework15 Nov 2023 0 repositories listed
-
14 Nov 2023 0 repositories listed
-
MART: Improving LLM Safety with Multi-round Automatic Red-Teaming13 Nov 2023 0 repositories listed
-
Summon a Demon and Bind it: A Grounded Theory of LLM Red Teaming10 Nov 2023 0 repositories listed
-
LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B31 Oct 2023 0 repositories listed
-
Learning from Red Teaming: Gender Bias Provocation and Mitigation in Large Language Models17 Oct 2023 0 repositories listed
-
Can Language Models be Instructed to Protect Personal Information?3 Oct 2023 0 repositories listed
-
Low-Resource Languages Jailbreak GPT-43 Oct 2023 0 repositories listed
-
Red Teaming Generative AI/NLP, the BB84 quantum cryptography protocol and the NIST-approved Quantum-Resistant Cryptographic Algorithms17 Sep 2023 0 repositories listed
-
The Promise and Peril of Artificial Intelligence -- Violet Teaming Offers a Balanced Path Forward28 Aug 2023 0 repositories listed
-
FLIRT: Feedback Loop In-context Red Teaming8 Aug 2023 0 repositories listed
-
11 Jul 2023 0 repositories listed
-
Seeing Seeds Beyond Weeds: Green Teaming Generative AI for Beneficial Uses30 May 2023 0 repositories listed
-
Personalisation within bounds: A risk taxonomy and policy framework for the alignment of large language models with personalised feedback9 Mar 2023 0 repositories listed
-
Red teaming ChatGPT via Jailbreaking: Bias, Robustness, Reliability and Toxicity30 Jan 2023 0 repositories listed
-
Can Large Language Models Change User Preference Adversarially?5 Jan 2023 0 repositories listed
-
Red-Teaming the Stable Diffusion Safety Filter3 Oct 2022 0 repositories listed
-
CTI4AI: Threat Intelligence Generation and Sharing after Red Teaming AI Models16 Aug 2022 0 repositories listed
-
Automating Privilege Escalation with Deep Reinforcement Learning4 Oct 2021 0 repositories listed
-
A Multi-Disciplinary Review of Knowledge Acquisition Methods: From Human to Autonomous Eliciting Agents27 Feb 2018 0 repositories listed
-
Computational Red Teaming in a Sudoku Solving Context: Neural Network Based Skill Representation and Acquisition27 Feb 2018 0 repositories listed
-
Shaping Influence and Influencing Shaping: A Computational Red Teaming Trust-based Swarm Intelligence Model26 Feb 2018 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.