Methods › Computer Vision › Generative Models › Sparse Autoencoder
Sparse Autoencoder
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
A Sparse Autoencoder is a type of autoencoder that employs sparsity to achieve an information bottleneck. Specifically the loss function is constructed so that activations are penalized within a layer. The sparsity constraint can be imposed with L1 regularization or a KL divergence between expected average neuron activation to an ideal distribution p.
Image: Jeff Jordan. Read his blog post (click) for a detailed summary of autoencoders.
Papers archive 2025-07-28
30 shown of 57, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Bridging Compositional and Distributional Semantics: A Survey on Latent Semantic Geometry via AutoEncoder 25 Jun 2025 · 0 repositories · arXiv:2506.20083
-
CWGAN-GP Augmented CAE for Jamming Detection in 5G-NR in Non-IID Datasets 18 Jun 2025 · 0 repositories · arXiv:2506.15075
-
Resa: Transparent Reasoning Models via SAEs 11 Jun 2025 · 1 repository · arXiv:2506.09967
-
Model Unlearning via Sparse Autoencoder Subspace Guided Projections 30 May 2025 · 0 repositories · arXiv:2505.24428
-
SAE-FiRE: Enhancing Earnings Surprise Predictions Through Sparse Autoencoder Feature Selection 20 May 2025 · 0 repositories · arXiv:2505.14420
-
Interpretable Risk Mitigation in LLM Agent Systems 15 May 2025 · 1 repository · arXiv:2505.10670
-
Are Sparse Autoencoders Useful for Java Function Bug Detection? 15 May 2025 · 1 repository · arXiv:2505.10375
-
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders 12 May 2025 · 0 repositories · arXiv:2505.08080Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)
-
Decoding Futures Price Dynamics: A Regularized Sparse Autoencoder for Interpretable Multi-Horizon Forecasting and Factor Discovery 11 May 2025 · 0 repositories · arXiv:2505.06795
-
Geospatial Mechanistic Interpretability of Large Language Models 6 May 2025 · 1 repository · arXiv:2505.03368
-
FineScope : Precision Pruning for Domain-Specialized Large Language Models Using SAE-Guided Self-Data Cultivation 1 May 2025 · 0 repositories · arXiv:2505.00624
-
Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition 29 Apr 2025 · 1 repository · arXiv:2504.20938
-
Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video 28 Apr 2025 · 2 repositories · arXiv:2504.19475Syntology ran 2 of 16 samples · 14 unverified
-
A real-time anomaly detection method for robots based on a flexible and sparse latent space 15 Apr 2025 · 1 repository · arXiv:2504.11170
-
Dissecting and Mitigating Diffusion Bias via Mechanistic Interpretability 26 Mar 2025 · 1 repository · arXiv:2503.20483Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)
-
Sparse Autoencoder as a Zero-Shot Classifier for Concept Erasing in Text-to-Image Diffusion Models 12 Mar 2025 · 1 repository · arXiv:2503.09446
-
Route Sparse Autoencoder to Interpret Large Language Models 11 Mar 2025 · 1 repository · arXiv:2503.08200Syntology ran 1 of 2 samples · 1 unverified
-
Self-Regularization with Latent Space Explanations for Controllable LLM-based Classification 19 Feb 2025 · 0 repositories · arXiv:2502.14133
-
LLM Pretraining with Continuous Concepts 12 Feb 2025 · 0 repositories · arXiv:2502.08524
-
Sparse Autoencoders for Hypothesis Generation 5 Feb 2025 · 1 repository · arXiv:2502.04382Syntology ran 2 of 2 samples · 0 unverified
-
LF-Steering: Latent Feature Activation Steering for Enhancing Semantic Consistency in Large Language Models 19 Jan 2025 · 0 repositories · arXiv:2501.11036
-
Steering Large Language Models with Feature Guided Activation Additions 17 Jan 2025 · 0 repositories · arXiv:2501.09929
-
Causal Graph Guided Steering of LLM Values via Prompts and Sparse Autoencoders 31 Dec 2024 · 0 repositories · arXiv:2501.00581
-
Sparse autoencoders reveal selective remapping of visual concepts during adaptation 6 Dec 2024 · 1 repository · arXiv:2412.05276Syntology ran 0 of 6 samples · 6 unverified
-
VISTA: A Panoramic View of Neural Representations 3 Dec 2024 · 0 repositories · arXiv:2412.02412
-
Direct Preference Optimization Using Sparse Feature-Level Constraints 12 Nov 2024 · 0 repositories · arXiv:2411.07618
-
Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders 27 Oct 2024 · 1 repository · arXiv:2410.20526Syntology ran 0 of 4 samples · 4 unverified · 4 pointer-only (licence)
-
Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs 16 Oct 2024 · 0 repositories · arXiv:2410.12555
-
Residual Stream Analysis with Multi-Layer SAEs 6 Sep 2024 · 1 repository · arXiv:2409.04185Syntology ran 3 of 3 samples · 0 unverified
-
Learning biologically relevant features in a pathology foundation model using sparse autoencoders 15 Jul 2024 · 0 repositories · arXiv:2407.10785
Tasks archive 2025-07-28
20 shown of 78 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
| Task | Papers |
|---|---|
| Classification | 4 |
| Denoising | 4 |
| General Classification | 4 |
| Dictionary Learning | 3 |
| Language Modeling | 3 |
| Language Modelling | 3 |
| Representation Learning | 3 |
| regression | 3 |
| Anomaly Detection | 2 |
| Clustering | 2 |
| Decoder | 2 |
| Diagnostic | 2 |
| Dimensionality Reduction | 2 |
| EEG | 2 |
| Electroencephalogram (EEG) | 2 |
| Generative Adversarial Network | 2 |
| Image Classification | 2 |
| Large Language Model | 2 |
| Small Data Image Classification | 2 |
| Transfer Learning | 2 |
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections