Browse State-of-the-Art › Safety Alignment › Papers, page 3
Safety Alignment
Papers archive 2025-07-28
archive papers tagged: 288 · with a code link: 134 · where Syntology ran a sample: 73 (59 with a run with no instrument failure, 14 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (73 of 288 tagged: 59 with a run with no instrument failure, 14 where every run was a failure of Syntology's instrument)
Page 3 of 3: papers 201 to 288 of 288, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Llama-3.1-Sherkala-8B-Chat: An Open Large Language Model for Kazakh3 Mar 2025 0 repositories listed
-
FC-Attack: Jailbreaking Multimodal Large Language Models via Auto-Generated Flowcharts28 Feb 2025 0 repositories listed
-
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence24 Feb 2025 0 repositories listed
-
Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignment21 Feb 2025 0 repositories listed
-
C3AI: Crafting and Evaluating Constitutions for Constitutional AI21 Feb 2025 0 repositories listed
-
Why Safeguarded Ships Run Aground? Aligned Large Language Models' Safety Mechanisms Tend to Be Anchored in The Template Region19 Feb 2025 0 repositories listed
-
Understanding and Rectifying Safety Perception Distortion in VLMs18 Feb 2025 0 repositories listed
-
DELMAN: Dynamic Defense Against Large Language Model Jailbreaking with Model Editing17 Feb 2025 0 repositories listed
-
Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models17 Feb 2025 0 repositories listed
-
VLM-Guard: Safeguarding Vision-Language Models via Fulfilling Safety Alignment Gap14 Feb 2025 0 repositories listed
-
Trustworthy AI: Safety, Bias, and Privacy -- A Survey11 Feb 2025 0 repositories listed
-
AI Alignment at Your Discretion10 Feb 2025 0 repositories listed
-
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions8 Feb 2025 0 repositories listed
-
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing4 Feb 2025 0 repositories listed
-
Internal Activation as the Polar Star for Steering Unsafe LLM Behavior3 Feb 2025 0 repositories listed
-
The dark deep side of DeepSeek: Fine-tuning attacks against the safety alignment of CoT-enabled models3 Feb 2025 0 repositories listed
-
Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning31 Jan 2025 0 repositories listed
-
Towards Safe AI Clinicians: A Comprehensive Study on Large Language Model Jailbreaking in Healthcare27 Jan 2025 0 repositories listed
-
Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models23 Jan 2025 0 repositories listed
-
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models7 Jan 2025 0 repositories listed
-
SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation3 Jan 2025 0 repositories listed
-
SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Models1 Jan 2025 0 repositories listed
-
No Free Lunch for Defending Against Prefilling Attack by In-Context Learning13 Dec 2024 0 repositories listed
-
SafetyDPO: Scalable Safety Alignment for Text-to-Image Generation13 Dec 2024 0 repositories listed
-
Model-Editing-Based Jailbreak against Safety-aligned Large Language Models11 Dec 2024 0 repositories listed
-
Na'vi or Knave: Jailbreaking Language Models via Metaphorical Avatars10 Dec 2024 0 repositories listed
-
Safety Alignment Backfires: Preventing the Re-emergence of Suppressed Concepts in Fine-tuned Text-to-Image Diffusion Models30 Nov 2024 0 repositories listed
-
PEFT-as-an-Attack! Jailbreaking Language Models during Federated Parameter-Efficient Fine-Tuning28 Nov 2024 0 repositories listed
-
Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models27 Nov 2024 0 repositories listed
-
Ensuring Safety and Trust: Analyzing the Risks of Large Language Models in Medicine20 Nov 2024 0 repositories listed
-
PSA-VLM: Enhancing Vision-Language Model Safety through Progressive Concept-Bottleneck-Driven Alignment18 Nov 2024 0 repositories listed
-
Playing Language Game with LLMs Leads to Jailbreaking16 Nov 2024 0 repositories listed
-
Unfair Alignment: Examining Safety Alignment Across Vision Encoder Layers in Vision-Language Models6 Nov 2024 0 repositories listed
-
Code-Switching Curriculum Learning for Multilingual Transfer in LLMs4 Nov 2024 0 repositories listed
-
Smaller Large Language Models Can Do Moral Self-Correction30 Oct 2024 0 repositories listed
-
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks23 Oct 2024 0 repositories listed
-
SPIN: Self-Supervised Prompt INjection17 Oct 2024 0 repositories listed
-
Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements11 Oct 2024 0 repositories listed
-
Unraveling and Mitigating Safety Alignment Degradation of Vision-Language Models11 Oct 2024 0 repositories listed
-
Superficial Safety Alignment Hypothesis7 Oct 2024 0 repositories listed
-
Toxic Subword Pruning for Dialogue Response Generation on Large Language Models5 Oct 2024 0 repositories listed
-
LLM Safeguard is a Double-Edged Sword: Exploiting False Positives for Denial-of-Service Attacks3 Oct 2024 0 repositories listed
-
SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks2 Oct 2024 0 repositories listed
-
Towards Inference-time Category-wise Safety Steering for Large Language Models2 Oct 2024 0 repositories listed
-
Backtracking Improves Generation Safety22 Sep 2024 0 repositories listed
-
PathSeeker: Exploring LLM Security Vulnerabilities with a Reinforcement Learning-Based Jailbreak Approach21 Sep 2024 0 repositories listed
-
Mitigating Unsafe Feedback with Learning Constraints19 Sep 2024 0 repositories listed
-
Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation20 Aug 2024 0 repositories listed
-
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming14 Aug 2024 0 repositories listed
-
EnJa: Ensemble Jailbreak on Large Language Models7 Aug 2024 0 repositories listed
-
Can Large Language Models Automatically Jailbreak GPT-4V?23 Jul 2024 0 repositories listed
-
Failures to Find Transferable Image Jailbreaks Between Vision-Language Models21 Jul 2024 0 repositories listed
-
Multilingual Blending: LLM Safety Alignment Evaluation with Language Mixture10 Jul 2024 0 repositories listed
-
Jailbreak Attacks and Defenses Against Large Language Models: A Survey5 Jul 2024 0 repositories listed
-
LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models3 Jul 2024 0 repositories listed
-
The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm26 Jun 2024 0 repositories listed
-
Adversarial Contrastive Decoding: Boosting Safety Alignment of Large Language Models via Opposite Prompt Optimization24 Jun 2024 0 repositories listed
-
20 Jun 2024 0 repositories listed Syntology 14 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 14 harvested samples) · 1 pointer-only (licence)
-
Model Merging and Safety Alignment: One Bad Model Spoils the Bunch20 Jun 2024 0 repositories listed
-
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference20 Jun 2024 0 repositories listed
-
Emerging Safety Attack and Defense in Federated Instruction Tuning of Large Language Models15 Jun 2024 0 repositories listed
-
Mimicking User Data: On Mitigating Fine-Tuning Risks in Closed Large Language Models12 Jun 2024 0 repositories listed
-
SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner8 Jun 2024 0 repositories listed
-
On the Intrinsic Self-Correction Capability of LLMs: Uncertainty and Latent Concept4 Jun 2024 0 repositories listed
-
Enhancing Jailbreak Attack Against Large Language Models through Silent Tokens31 May 2024 0 repositories listed
-
Cross-Modal Safety Alignment: Is textual unlearning all you need?27 May 2024 0 repositories listed
-
No Two Devils Alike: Unveiling Distinct Mechanisms of Fine-tuning Attacks25 May 2024 0 repositories listed
-
Robustifying Safety-Aligned Large Language Models through Clean Data Curation24 May 2024 0 repositories listed
-
Safety Alignment for Vision Language Models22 May 2024 0 repositories listed
-
Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching22 May 2024 0 repositories listed
-
WordGame: Efficient & Effective LLM Jailbreak via Simultaneous Obfuscation in Query and Response22 May 2024 0 repositories listed
-
CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues4 Apr 2024 0 repositories listed
-
Learn to Disguise: Avoid Refusal Responses in LLM's Defense via a Multi-agent Attacker-Disguiser Game3 Apr 2024 0 repositories listed
-
Enhancing Jailbreak Attacks with Diversity Guidance1 Mar 2024 0 repositories listed
-
LLMs Can Defend Themselves Against Jailbreaking in a Practical Manner: A Vision Paper24 Feb 2024 0 repositories listed
-
Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement23 Feb 2024 0 repositories listed
-
Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications7 Feb 2024 0 repositories listed
-
Cognitive Overload: Jailbreaking Large Language Models with Overloaded Logical Thinking16 Nov 2023 0 repositories listed
-
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models16 Nov 2023 0 repositories listed
-
MART: Improving LLM Safety with Multi-round Automatic Red-Teaming13 Nov 2023 0 repositories listed
-
LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B31 Oct 2023 0 repositories listed
-
Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks16 Oct 2023 0 repositories listed
-
Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations10 Oct 2023 0 repositories listed
-
Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models4 Oct 2023 0 repositories listed
-
Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models30 Aug 2023 0 repositories listed
-
Deceptive Alignment Monitoring20 Jul 2023 0 repositories listed
-
11 Jul 2023 0 repositories listed
-
Off-Policy Risk Assessment in Markov Decision Processes21 Sep 2022 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.