Browse State-of-the-Art › backdoor defense
backdoor defense
59 papers with code · 0 benchmarks · 2 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
2 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 59 papers with code (131 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
5 Feb 2022 3 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Recent studies have revealed that deep neural networks (DNNs) are vulnerable to backdoor attacks, where attackers embed hidden backdoors in the DNN model by poisoning a few training samples.
-
2 Dec 2021 3 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedHowever, designing a unified BA method that can be applied to various MIA systems is challenging due to the diversity of imaging modalities (e.
-
22 Feb 2025 2 repositories listed Syntology ran 5 of 10 samples · 5 unverified · 10 pointer-only (licence)Backdoor attacks on deep neural networks (DNNs) have emerged as a significant security threat, allowing adversaries to implant hidden malicious behaviors during the model training phase.
-
13 Oct 2024 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)We find that current safety purification methods are vulnerable to the rapid re-learning of backdoor behavior, even when further fine-tuning of purified models is performed using a very small number of poisoned samples.
-
25 May 2024 2 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 5 pointer-only (licence)Specifically, PDB leverages the home-field advantage of defenders by proactively injecting a defensive backdoor into the model during training.
-
28 Jul 2023 2 repositories listed Syntology ran 5 of 9 samples · 4 unverified · 9 pointer-only (licence)Deep neural networks (DNNs) are vulnerable to backdoor attack, which does not affect the network's performance on clean data but would manipulate the network behavior once a trigger pattern is added.
-
20 Jul 2023 2 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)By establishing the connection between backdoor risk and adversarial risk, we derive a novel upper bound for backdoor risk, which mainly captures the risk on the shared adversarial examples (SAEs) between the backdoored…
-
1 Jan 2021 2 repositories listedUnder this optimization framework, the trigger generator function will learn to manipulate the input with imperceptible noise to preserve the model performance on the clean data and maximize the attack success rate on…
-
20 Nov 2020 2 repositories listed Syntology ran 3 of 4 samples · 1 unverifiedNevertheless, there are few studies on defending against textual backdoor attacks.
-
7 Jul 2025 1 repository listedCGD utilizes a publicly accessible CLIP model to identify inputs that are likely to be clean or poisoned.
-
17 May 2025 1 repository listedIn particular, we divide the local model into a feature extractor and a classifier.
-
30 Apr 2025 1 repository listedTo conquer this challenge, we introduce a storage-update-based certification method, which dynamically adjusts each sample's certification region to improve certification performance.
-
22 Apr 2025 1 repository listedWe proceed to prove the feasibility of activating redundant neurons utilizing out-of-distribution (OOD) samples in centralized settings, and migrating to FL settings to propose a novel backdoor defense mechanism,…
-
28 Feb 2025 1 repository listedExisting defense strategies are well equipped to thwart such attacks through backdoor detection and trigger inversion because previous attack methods are constrained by limited input spaces and low-dimensional triggers.
-
10 Jan 2025 1 repository listedTo answer this question, we examine 12 common backdoor attacks that focus on input-space or feature-space stealthiness and 17 diverse representative defenses.
-
5 Jan 2025 1 repository listedThe BTU defense leverages these properties to identify aberrant embedding parameters and subsequently removes backdoor behaviors using a fine-grained unlearning technique.
-
3 Dec 2024 1 repository listedBackdoor attacks remain significant security threats to generative large language models (LLMs).
-
18 Nov 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverifiedRecent studies reveal that Large Language Models (LLMs) are susceptible to backdoor attacks, where adversaries embed hidden triggers that manipulate model responses.
-
17 Nov 2024 1 repository listedIn order to facilitate the research in multimodal backdoor, we introduce BackdoorMBTI, the first backdoor learning toolkit and benchmark designed for multimodal evaluation across three representative modalities from…
-
25 Oct 2024 1 repository listedSpecifically, EBYD first exposes the backdoor functionality in the backdoored model through a model preprocessing step called backdoor exposure, and then applies detection and removal methods to the exposed model to…
-
2 Oct 2024 1 repository listedTo achieve this objective, we ask: How to recover universal and hard backdoor triggers in GNNs?
-
9 Sep 2024 1 repository listedAdditionally, with the reversed trigger, we propose backdoor detection from the noise space, introducing the first backdoor input detection approach for diffusion models and a novel model detection algorithm that…
-
1 Sep 2024 1 repository listedIntuitively, the backdoor can be purified by re-optimizing the model to smoother minima.
-
28 Aug 2024 1 repository listedSubsequently, VFLIP conducts purification which removes the embeddings identified as malicious and reconstructs all the embeddings based on the remaining embeddings.
-
31 Jul 2024 1 repository listedWe evaluate our framework on hundreds of DMs that are attacked by three existing backdoor attack methods with a wide range of hyperparameter settings.
-
28 May 2024 1 repository listedSpecifically, our PUD has a progressive model purification scheme to jointly erase backdoors and enhance the model's adversarial robustness.
-
18 May 2024 1 repository listed Syntology ran 3 of 3 samples · 0 unverifiedBackdoor attacks pose an increasingly severe security threat to Deep Neural Networks (DNNs) during their development stage.
-
15 Mar 2024 1 repository listed Syntology ran 4 of 5 samples · 1 unverified · 5 pointer-only (licence)Based on this, we pose the backdoor data identification problem as a hierarchical data splitting optimization problem, leveraging a novel SPC-based loss function as the primary optimization objective.
-
17 Feb 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)In this work, we take the first step to investigate one of the typical safety threats, backdoor attack, to LLM-based agents.
-
4 Jan 2024 1 repository listed Syntology ran 5 of 9 samples · 4 unverified · 9 pointer-only (licence)Backdoor attack aims to deceive a victim model when facing backdoor instances while maintaining its performance on benign data.
Syntology lines on 13 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections