{"url":"/method/sparse-autoencoder","slug":"sparse-autoencoder","name":"Sparse Autoencoder","full_name":"Sparse Autoencoder","full_name_withheld":false,"description_markdown":"A **Sparse Autoencoder** is a type of autoencoder that employs sparsity to achieve an information bottleneck. Specifically the loss function is constructed so that activations are penalized within a layer. The sparsity constraint can be imposed with [L1 regularization](https://paperswithcode.com/method/l1-regularization) or a KL divergence between expected average neuron activation to an ideal distribution $p$.\r\n\r\nImage: [Jeff Jordan](https://www.jeremyjordan.me/autoencoders/). Read his blog post (click) for a detailed summary of autoencoders.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":null,"title":null,"url_on_a_paper_host":false},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Generative Models","url":"/methods/category/generative-models","pwc_aliases":[]}],"n_papers_tagged":57,"archive_num_papers":57,"papers_newest_first":[{"paper":null,"title":"Bridging Compositional and Distributional Semantics: A Survey on Latent Semantic Geometry via AutoEncoder","date":"2025-06-25","arxiv_id":"2506.20083","n_code_links":0,"syntology":null},{"paper":null,"title":"CWGAN-GP Augmented CAE for Jamming Detection in 5G-NR in Non-IID Datasets","date":"2025-06-18","arxiv_id":"2506.15075","n_code_links":0,"syntology":null},{"paper":"/paper/resa-transparent-reasoning-models-via-saes","title":"Resa: Transparent Reasoning Models via SAEs","date":"2025-06-11","arxiv_id":"2506.09967","n_code_links":1,"syntology":null},{"paper":null,"title":"Model Unlearning via Sparse Autoencoder Subspace Guided Projections","date":"2025-05-30","arxiv_id":"2505.24428","n_code_links":0,"syntology":null},{"paper":null,"title":"SAE-FiRE: Enhancing Earnings Surprise Predictions Through Sparse Autoencoder Feature Selection","date":"2025-05-20","arxiv_id":"2505.14420","n_code_links":0,"syntology":null},{"paper":"/paper/2505-10670","title":"Interpretable Risk Mitigation in LLM Agent Systems","date":"2025-05-15","arxiv_id":"2505.10670","n_code_links":1,"syntology":null},{"paper":"/paper/are-sparse-autoencoders-useful-for-java","title":"Are Sparse Autoencoders Useful for Java Function Bug Detection?","date":"2025-05-15","arxiv_id":"2505.10375","n_code_links":1,"syntology":null},{"paper":"/paper/beyond-input-activations-identifying","title":"Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders","date":"2025-05-12","arxiv_id":"2505.08080","n_code_links":0,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":1}},{"paper":null,"title":"Decoding Futures Price Dynamics: A Regularized Sparse Autoencoder for Interpretable Multi-Horizon Forecasting and Factor Discovery","date":"2025-05-11","arxiv_id":"2505.06795","n_code_links":0,"syntology":null},{"paper":"/paper/geospatial-mechanistic-interpretability-of","title":"Geospatial Mechanistic Interpretability of Large Language Models","date":"2025-05-06","arxiv_id":"2505.03368","n_code_links":1,"syntology":null},{"paper":null,"title":"FineScope : Precision Pruning for Domain-Specialized Large Language Models Using SAE-Guided Self-Data Cultivation","date":"2025-05-01","arxiv_id":"2505.00624","n_code_links":0,"syntology":null},{"paper":"/paper/towards-understanding-the-nature-of-attention","title":"Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition","date":"2025-04-29","arxiv_id":"2504.20938","n_code_links":1,"syntology":null},{"paper":"/paper/prisma-an-open-source-toolkit-for-mechanistic","title":"Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video","date":"2025-04-28","arxiv_id":"2504.19475","n_code_links":2,"syntology":{"ran":2,"of":16,"unverified":14,"pointer_only":0}},{"paper":"/paper/a-real-time-anomaly-detection-method-for","title":"A real-time anomaly detection method for robots based on a flexible and sparse latent space","date":"2025-04-15","arxiv_id":"2504.11170","n_code_links":1,"syntology":null},{"paper":"/paper/dissecting-and-mitigating-diffusion-bias-via","title":"Dissecting and Mitigating Diffusion Bias via Mechanistic Interpretability","date":"2025-03-26","arxiv_id":"2503.20483","n_code_links":1,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":1}},{"paper":"/paper/sparse-autoencoder-as-a-zero-shot-classifier","title":"Sparse Autoencoder as a Zero-Shot Classifier for Concept Erasing in Text-to-Image Diffusion Models","date":"2025-03-12","arxiv_id":"2503.09446","n_code_links":1,"syntology":null},{"paper":"/paper/route-sparse-autoencoder-to-interpret-large","title":"Route Sparse Autoencoder to Interpret Large Language Models","date":"2025-03-11","arxiv_id":"2503.08200","n_code_links":1,"syntology":{"ran":1,"of":2,"unverified":1,"pointer_only":0}},{"paper":null,"title":"Self-Regularization with Latent Space Explanations for Controllable LLM-based Classification","date":"2025-02-19","arxiv_id":"2502.14133","n_code_links":0,"syntology":null},{"paper":null,"title":"LLM Pretraining with Continuous Concepts","date":"2025-02-12","arxiv_id":"2502.08524","n_code_links":0,"syntology":null},{"paper":"/paper/sparse-autoencoders-for-hypothesis-generation","title":"Sparse Autoencoders for Hypothesis Generation","date":"2025-02-05","arxiv_id":"2502.04382","n_code_links":1,"syntology":{"ran":2,"of":2,"unverified":0,"pointer_only":0}},{"paper":null,"title":"LF-Steering: Latent Feature Activation Steering for Enhancing Semantic Consistency in Large Language Models","date":"2025-01-19","arxiv_id":"2501.11036","n_code_links":0,"syntology":null},{"paper":null,"title":"Steering Large Language Models with Feature Guided Activation Additions","date":"2025-01-17","arxiv_id":"2501.09929","n_code_links":0,"syntology":null},{"paper":null,"title":"Causal Graph Guided Steering of LLM Values via Prompts and Sparse Autoencoders","date":"2024-12-31","arxiv_id":"2501.00581","n_code_links":0,"syntology":null},{"paper":"/paper/sparse-autoencoders-reveal-selective","title":"Sparse autoencoders reveal selective remapping of visual concepts during adaptation","date":"2024-12-06","arxiv_id":"2412.05276","n_code_links":1,"syntology":{"ran":0,"of":6,"unverified":6,"pointer_only":0}},{"paper":null,"title":"VISTA: A Panoramic View of Neural Representations","date":"2024-12-03","arxiv_id":"2412.02412","n_code_links":0,"syntology":null},{"paper":null,"title":"Direct Preference Optimization Using Sparse Feature-Level Constraints","date":"2024-11-12","arxiv_id":"2411.07618","n_code_links":0,"syntology":null},{"paper":"/paper/llama-scope-extracting-millions-of-features","title":"Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders","date":"2024-10-27","arxiv_id":"2410.20526","n_code_links":1,"syntology":{"ran":0,"of":4,"unverified":4,"pointer_only":4}},{"paper":null,"title":"Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs","date":"2024-10-16","arxiv_id":"2410.12555","n_code_links":0,"syntology":null},{"paper":"/paper/residual-stream-analysis-with-multi-layer","title":"Residual Stream Analysis with Multi-Layer SAEs","date":"2024-09-06","arxiv_id":"2409.04185","n_code_links":1,"syntology":{"ran":3,"of":3,"unverified":0,"pointer_only":0}},{"paper":null,"title":"Learning biologically relevant features in a pathology foundation model using sparse autoencoders","date":"2024-07-15","arxiv_id":"2407.10785","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/classification-1","name":"Classification","papers":4},{"task":"/task/denoising","name":"Denoising","papers":4},{"task":"/task/classification","name":"General Classification","papers":4},{"task":"/task/dictionary-learning","name":"Dictionary Learning","papers":3},{"task":"/task/language-modeling","name":"Language Modeling","papers":3},{"task":"/task/language-modelling","name":"Language Modelling","papers":3},{"task":"/task/representation-learning","name":"Representation Learning","papers":3},{"task":"/task/regression-1","name":"regression","papers":3},{"task":"/task/anomaly-detection","name":"Anomaly Detection","papers":2},{"task":"/task/clustering","name":"Clustering","papers":2},{"task":"/task/decoder","name":"Decoder","papers":2},{"task":"/task/diagnostic","name":"Diagnostic","papers":2},{"task":"/task/dimensionality-reduction","name":"Dimensionality Reduction","papers":2},{"task":"/task/eeg-1","name":"EEG","papers":2},{"task":"/task/eeg","name":"Electroencephalogram (EEG)","papers":2},{"task":null,"name":"Generative Adversarial Network","papers":2},{"task":"/task/image-classification","name":"Image Classification","papers":2},{"task":"/task/large-language-model","name":"Large Language Model","papers":2},{"task":"/task/small-data","name":"Small Data Image Classification","papers":2},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":2}],"tasks_shown":20,"n_tasks":78,"usage_by_year":[{"year":"2011","papers":1},{"year":"2015","papers":2},{"year":"2016","papers":2},{"year":"2017","papers":1},{"year":"2018","papers":3},{"year":"2019","papers":5},{"year":"2020","papers":2},{"year":"2021","papers":1},{"year":"2022","papers":2},{"year":"2023","papers":4},{"year":"2024","papers":12},{"year":"2025","papers":22}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/sparse-autoencoder"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}