Methods › Natural Language Processing › Inference Extrapolation › Patching
Activation Patching
Patching
Introduced by Asma Ghandeharioun et al. in Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Activation patching studies the model's computation by altering its latent representations, the token embeddings in transformer-based language models, during the inference process
Papers archive 2025-07-28
30 shown of 102, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Data Augmentation in Time Series Forecasting through Inverted Framework 15 Jul 2025 · 0 repositories · arXiv:2507.11439
-
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers 12 Jul 2025 · 0 repositories · arXiv:2507.09406
-
Unpatchable Vulnerabilities in Windows 10/11: Security Report 2025 10 Jul 2025 · 0 repositories
-
Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models 25 Jun 2025 · 1 repository · arXiv:2506.20251Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)
-
SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks 13 Jun 2025 · 1 repository · arXiv:2506.11791
-
Time Series Representations for Classification Lie Hidden in Pretrained Vision Transformers 10 Jun 2025 · 0 repositories · arXiv:2506.08641
-
Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models 9 Jun 2025 · 1 repository · arXiv:2506.07468Syntology ran 6 of 17 samples · 11 unverified · 1 pointer-only (licence)
-
A Multi-Dataset Evaluation of Models for Automated Vulnerability Repair 5 Jun 2025 · 0 repositories · arXiv:2506.04987
-
Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation 29 May 2025 · 0 repositories · arXiv:2505.23701
-
Foundation Model for Wireless Technology Recognition Using IQ Timeseries 26 May 2025 · 0 repositories · arXiv:2505.19390
-
From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance 26 May 2025 · 1 repository · arXiv:2505.20229
-
Co-PatcheR: Collaborative Software Patching with Component(s)-specific Small Reasoning Models 25 May 2025 · 0 repositories · arXiv:2505.18955
-
Know the Ropes: A Heuristic Strategy for LLM-based Multi-Agent System Design 22 May 2025 · 0 repositories · arXiv:2505.16979
-
Internal Chain-of-Thought: Empirical Evidence for Layer-wise Subtask Scheduling in LLMs 20 May 2025 · 1 repository · arXiv:2505.14530
-
SPIRIT: Patching Speech Language Models against Jailbreak Attacks 18 May 2025 · 0 repositories · arXiv:2505.13541
-
SPAT: Sensitivity-based Multihead-attention Pruning on Time Series Forecasting Models 13 May 2025 · 0 repositories · arXiv:2505.08768
-
Unpacking Robustness in Inflectional Languages: Adversarial Evaluation and Mechanistic Insights 8 May 2025 · 0 repositories · arXiv:2505.07856
-
Interpreting Multilingual and Document-Length Sensitive Relevance Computations in Neural Retrieval Models through Axiomatic Causal Interventions 4 May 2025 · 1 repository · arXiv:2505.02154
-
The Illusion of Role Separation: Hidden Shortcuts in LLM Role Learning (and How to Fix Them) 1 May 2025 · 0 repositories · arXiv:2505.00626
-
HSE: A plug-and-play module for unified fault diagnosis foundation models 26 Apr 2025 · 1 repository
-
MIB: A Mechanistic Interpretability Benchmark 17 Apr 2025 · 1 repository · arXiv:2504.13151Syntology ran 3 of 3 samples · 0 unverified
-
How do Large Language Models Understand Relevance? A Mechanistic Interpretability Perspective 10 Apr 2025 · 1 repository · arXiv:2504.07898
-
Localized Definitions and Distributed Reasoning: A Proof-of-Concept Mechanistic Interpretability Study via Activation Patching 3 Apr 2025 · 1 repository · arXiv:2504.02976
-
Reverse-Engineering the Retrieval Process in GenIR Models 25 Mar 2025 · 1 repository · arXiv:2503.19715
-
Sentinel: Multi-Patch Transformer with Temporal and Channel Attention for Time Series Forecasting 22 Mar 2025 · 0 repositories · arXiv:2503.17658
-
A Semantic-based Optimization Approach for Repairing LLMs: Case Study on Code Generation 17 Mar 2025 · 0 repositories · arXiv:2503.12899
-
TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research 17 Mar 2025 · 0 repositories · arXiv:2503.12730
-
Rethinking Lanes and Points in Complex Scenarios for Monocular 3D Lane Detection 8 Mar 2025 · 0 repositories · arXiv:2503.06237
-
TimeFound: A Foundation Model for Time Series Forecasting 6 Mar 2025 · 0 repositories · arXiv:2503.04118
-
Superscopes: Amplifying Internal Feature Representations for Language Model Interpretation 3 Mar 2025 · 1 repository · arXiv:2503.02078
Tasks archive 2025-07-28
20 shown of 118 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections