Browse State-of-the-Art › Safety Alignment
Safety Alignment
134 papers with code · 0 benchmarks · 1 dataset archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 134 papers with code (288 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
13 Feb 2025 3 repositories listedWe then measure how different directions promote or suppress the dominant direction, showing the important role of secondary directions in shaping the model's refusal representation.
-
18 Aug 2024 3 repositories listed Syntology ran 3 of 5 samples · 2 unverifiedTo this end, we propose Antidote, a post-fine-tuning stage solution, which remains \textbf{\textit{agnostic to the training hyper-parameters in the fine-tuning stage}}.
-
14 Oct 2024 2 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)This study exposes the safety vulnerabilities of Large Language Models (LLMs) in multi-turn interactions, where malicious users can obscure harmful intents across several queries.
-
26 Sep 2024 2 repositories listed Syntology ran 6 of 9 samples · 3 unverifiedTo clear up concern, this paper provide a comprehensive overview to three aspects of harmful fine-tuning: attacks setting, defense design and evaluation methodology.
-
12 Mar 2024 2 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedThe rapid advancement of Large Language Models (LLMs) has brought about remarkable generative capabilities but also raised concerns about their potential misuse.
-
9 Nov 2023 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Large Vision-Language Models (LVLMs) signify a groundbreaking paradigm shift within the Artificial Intelligence (AI) community, extending beyond the capabilities of Large Language Models (LLMs) by assimilating…
-
5 Oct 2023 2 repositories listedA single language model, even when aligned with labelers through reinforcement learning from human feedback (RLHF), may not suit all human preferences.
-
18 Aug 2023 2 repositories listedIn this work, we propose a new safety evaluation benchmark RED-EVAL that carries out red-teaming.
-
15 Jul 2025 1 repository listed Syntology ran 11 of 17 samples · 6 unverifiedDiffusion-based large language models (dLLMs) have recently emerged as a powerful alternative to autoregressive LLMs, offering faster inference and greater interactivity via parallel decoding and bidirectional modeling.
-
21 Jun 2025 1 repository listed Syntology ran 0 of 8 samples · 8 unverified · 8 pointer-only (licence)Fine-tuning Large Language Models (LLMs) with Low-Rank Adaptation (LoRA) enhances adaptability while reducing computational costs.
-
19 Jun 2025 1 repository listedBackdoor unalignment attacks against Large Language Models (LLMs) enable the stealthy compromise of safety alignment using a hidden trigger while evading normal safety auditing.
-
19 Jun 2025 1 repository listedConsequently, small shifts in hidden activations can re-trigger harmful behaviors embedded in the latent space.
-
16 Jun 2025 1 repository listedLarge language models (LLMs) have shown strong performance across natural language tasks, but remain vulnerable to backdoor attacks.
-
12 Jun 2025 1 repository listedWe show that a carefully prompt engineered lightweight monitor achieves a 93% defense success rate, beating reasoning models like o3 mini as a monitor.
-
11 Jun 2025 1 repository listedExtensive experiments across five benchmarks on two representative LVLMs demonstrate that DAVSP effectively resists malicious queries while preserving benign input utility.
-
9 Jun 2025 1 repository listedLarge Language Models (LLMs) continue to exhibit vulnerabilities despite deliberate safety alignment efforts, posing significant risks to users and society.
-
9 Jun 2025 1 repository listed Syntology ran 6 of 17 samples · 11 unverified · 1 pointer-only (licence)Conventional language model (LM) safety alignment relies on a reactive, disjoint procedure: attackers exploit a static model, followed by defensive fine-tuning to patch exposed vulnerabilities.
-
3 Jun 2025 1 repository listedThe inherent risk of generating harmful and unsafe content by Large Language Models (LLMs), has highlighted the need for their safety alignment.
-
3 Jun 2025 1 repository listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Finetuning is a critical step for adapting large language models (LLMs) to domain-specific downstream tasks.
-
30 May 2025 1 repository listedLarge Language Models (LLMs) excel in various natural language processing tasks but remain vulnerable to generating harmful content or being exploited for malicious purposes.
-
29 May 2025 1 repository listedThe acquisition of agentic capabilities has transformed LLMs from "knowledge providers" to "action executors", a trend that while expanding LLMs' capability boundaries, significantly increases their susceptibility to…
-
27 May 2025 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedText-to-Image (T2I) models have achieved remarkable success in generating visual content from text inputs.
-
26 May 2025 1 repository listedLLMs have made impressive progress, but their growing capabilities also expose them to highly flexible jailbreaking attacks designed to bypass safety alignment.
-
26 May 2025 1 repository listedDespite the remarkable proficiency of \textit{Large Reasoning Models} (LRMs) in handling complex reasoning tasks, their reliability in safety-critical scenarios remains uncertain.
-
26 May 2025 1 repository listedIn this paper, we introduce the concept of safety calibration, which systematically addresses both undersafety and oversafety.
-
23 May 2025 1 repository listedInvestigating these attacks is crucial for uncovering model vulnerabilities.
-
22 May 2025 1 repository listedThe significant progress of large language models (LLMs) has led to remarkable achievements across numerous applications.
-
22 May 2025 1 repository listedLarge language models (LLMs) are considered valuable Intellectual Properties (IP) for legitimate owners due to the enormous computational cost of training.
-
22 May 2025 1 repository listed Syntology ran 2 of 3 samples · 1 unverifiedLarge language models (LLMs) have become increasingly central to AI applications worldwide, necessitating robust multilingual safety alignment to ensure secure deployment across diverse linguistic contexts.
-
22 May 2025 1 repository listedIn the adversarial iterative optimization stage, the red-team model and the target model continuously improve their respective capabilities in interaction.
Syntology lines on 11 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections