Browse State-of-the-Art › Model Editing
Model Editing
107 papers with code · 0 benchmarks · 1 dataset archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 107 papers with code (193 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
22 May 2023 4 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedOur objective is to provide valuable insights into the effectiveness and feasibility of each editing technique, thereby assisting the community in making informed decisions on the selection of the most appropriate…
-
10 Feb 2022 4 repositories listed Syntology ran 4 of 15 samples · 11 unverifiedTo test our hypothesis that these computations correspond to factual association recall, we modify feed-forward weights to update specific factual associations using Rank-One Model Editing (ROME).
-
12 Mar 2024 3 repositories listed Syntology ran 1 of 10 samples · 9 unverifiedInterventions on model-internal states are fundamental operations in many areas of AI, including model editing, steering, robustness, and interpretability.
-
15 Sep 2023 3 repositories listed Syntology ran 4 of 11 samples · 7 unverified · 5 pointer-only (licence)One hypothesised cause of polysemanticity is \textit{superposition}, where neural networks represent more features than they have neurons by assigning features to an overcomplete set of directions in activation space,…
-
21 Oct 2021 3 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedTo enable easy post-hoc editing at scale, we propose Model Editor Networks using Gradient Decomposition (MEND), a collection of small auxiliary editing networks that use a single desired input-output pair to make fast,…
-
16 Feb 2025 2 repositories listed Syntology ran 5 of 8 samples · 3 unverifiedOur analysis provides a fundamental reexamination of both the real-world applicability of existing model editing methods and their evaluation practices, and establishes a rigorous evaluation framework with key insights…
-
5 Oct 2024 2 repositories listed Syntology ran 7 of 7 samples · 0 unverified · 7 pointer-only (licence)This work explores sequential model editing in large language models (LLMs), a critical task that involves modifying internal knowledge within LLMs continuously through multi-round editing, each incorporating updates or…
-
3 Oct 2024 2 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 1 pointer-only (licence)To address this, we introduce AlphaEdit, a novel solution that projects perturbation onto the null space of the preserved knowledge before applying it to the parameters.
-
21 Sep 2024 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)We find arithmetic ability resides within a limited number of attention heads, with each head specializing in distinct operations.
-
22 Jul 2024 2 repositories listed Syntology ran 5 of 11 samples · 6 unverifiedFor example, the LLM red-teaming literature has produced a wide variety of 'jailbreaking' techniques to elicit harmful text from models that were fine-tuned to be harmless.
-
22 May 2024 2 repositories listed Syntology ran 7 of 12 samples · 5 unverifiedFurthermore, these tuning-based methods require large-scale preference data for training and are susceptible to noisy preference data.
-
2 Jan 2024 2 repositories listedIn this paper, we first define the knowledge editing problem and then provide a comprehensive review of cutting-edge approaches.
-
14 Mar 2023 2 repositories listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)Our Text-to-Image Model Editing method, TIME for short, receives a pair of inputs: a "source" under-specified prompt for which the model makes an implicit assumption (e.
-
30 Jun 2022 2 repositories listedMachine learning (ML) interpretability techniques can reveal undesirable patterns in data that models exploit to make predictions--potentially causing harms once deployed.
-
25 Jun 2025 1 repository listedTo efficiently steer the ethical behavior of agents, we frame agent behavior steering as a model editing task, which we term Behavior Editing.
-
16 Jun 2025 1 repository listedLarge language models (LLMs) have shown strong performance across natural language tasks, but remain vulnerable to backdoor attacks.
-
30 May 2025 1 repository listedOriginally, dropout was seen as a breakthrough regularization technique that reduced overfitting and improved performance in almost all applications of deep learning by reducing overfitting.
-
21 May 2025 1 repository listedLarge language models require iterative updates to address challenges such as knowledge conflicts and outdated information (e.
-
21 May 2025 1 repository listedTo tackle this, we model the sequential editing as a constrained stochastic programming.
-
20 May 2025 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Lifelong learning enables large language models (LLMs) to adapt to evolving information by continually updating their internal knowledge.
-
17 May 2025 1 repository listedTo address this challenge, we propose a method based on few-shot orthogonal alignment, which aligns task vectors to the parameter space of a differently pre-trained target model.
-
17 May 2025 1 repository listed Syntology ran 3 of 3 samples · 0 unverifiedModel editing techniques are essential for efficiently updating knowledge in large language models (LLMs).
-
3 Apr 2025 1 repository listedHowever, existing methods rely on network linearization to derive task vectors, leading to computational bottlenecks during training and inference.
-
3 Apr 2025 1 repository listedThe stronger clean logit difference for definitional questions further supports this localized representation.
-
29 Mar 2025 1 repository listedLastly, Rank One Model Editing (ROME) is used to remove the password information from the model, resulting in the number of passwords recovered going from 37 to 0.
-
11 Mar 2025 1 repository listedPrevious studies have established that language models manifest stereotyped biases.
-
10 Mar 2025 1 repository listed Syntology ran 1 of 3 samples · 2 unverifiedErasing concepts from large-scale text-to-image (T2I) diffusion models has become increasingly crucial due to the growing concerns over copyright infringement, offensive content, and privacy violations.
-
20 Feb 2025 1 repository listedLarge language models (LLMs) often retain outdated or incorrect information from pre-training, which undermines their reliability.
-
9 Feb 2025 1 repository listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)Our findings underscore the effectiveness, stealthiness, and explainability of JailbreakEdit, emphasizing the need for more advanced defense mechanisms in LLMs.
-
9 Feb 2025 1 repository listedLarge language models (LLMs) acquire information from pre-training corpora, but their stored knowledge can become inaccurate or outdated over time.
Syntology lines on 16 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections