{"url":"/method/mamba","slug":"mamba","name":"Mamba","full_name":"Mamba: Linear-Time Sequence Modeling with Selective State Spaces","full_name_withheld":false,"description_markdown":"Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module. Many subquadratic-time architectures such as linear attention, gated convolution and recurrent models, and structured state space models (SSMs) have been developed to address Transformers’ computational inefficiency on long sequences, but they have not performed as well as attention on important modalities such as language. We identify that a key weakness of such models is their inability to perform content-based reasoning, and make several improvements. First, simply letting the SSM parameters be functions of the input addresses their weakness with discrete modalities, allowing the model to selectively propagate or forget information along the sequence length dimension depending on the current token. Second, even though this change prevents the use of efficient convolutions, we design a hardware-aware parallel algorithm in recurrent mode. We integrate these selective SSMs into a simplified end-to-end neural network architecture without attention or even MLP blocks (Mamba). Mamba enjoys fast inference (5× higher throughput than Transformers) and linear scaling in sequence length, and its performance improves on real data up to million-length sequences. As a general sequence model backbone, Mamba achieves state-of-the-art performance across several modalities such as language, audio, and genomics. On language modeling, our Mamba-3B model outperforms Transformers of the same size and matches Transformers twice its size, both in pre-training and downstream evaluation.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Mamba: Linear-Time Sequence Modeling with Selective State Spaces","paper":"/paper/mamba-linear-time-sequence-modeling-with","first_author":"Albert Gu","n_authors":2,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/mamba-linear-time-sequence-modeling-with"},"source":{"url":"https://arxiv.org/abs/2312.00752v2","title":"Mamba: Linear-Time Sequence Modeling with Selective State Spaces","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/state-spaces/mamba","code_snippet_url_on_a_code_host":true,"categories":[],"n_papers_tagged":599,"archive_num_papers":599,"papers_newest_first":[{"paper":"/paper/differential-mamba","title":"Differential Mamba","date":"2025-07-08","arxiv_id":"2507.06204","n_code_links":1,"syntology":{"ran":2,"of":2,"unverified":0,"pointer_only":0}},{"paper":"/paper/langmamba-a-language-driven-mamba-framework","title":"LangMamba: A Language-driven Mamba Framework for Low-dose CT Denoising with Vision-language Models","date":"2025-07-08","arxiv_id":"2507.06140","n_code_links":1,"syntology":null},{"paper":"/paper/findrec-stein-guided-entropic-flow-for-multi","title":"FindRec: Stein-Guided Entropic Flow for Multi-Modal Sequential Recommendation","date":"2025-07-07","arxiv_id":"2507.04651","n_code_links":1,"syntology":null},{"paper":"/paper/mambafusion-height-fidelity-dense-global","title":"MambaFusion: Height-Fidelity Dense Global Fusion for Multi-modal 3D Object Detection","date":"2025-07-06","arxiv_id":"2507.04369","n_code_links":1,"syntology":{"ran":0,"of":10,"unverified":10,"pointer_only":0}},{"paper":"/paper/mvnet-hyperspectral-remote-sensing-image","title":"MVNet: Hyperspectral Remote Sensing Image Classification Based on Hybrid Mamba-Transformer Vision Backbone Architecture","date":"2025-07-06","arxiv_id":"2507.04409","n_code_links":1,"syntology":null},{"paper":"/paper/mamba-guided-boundary-prior-matters-a-new","title":"Mamba Guided Boundary Prior Matters: A New Perspective for Generalized Polyp Segmentation","date":"2025-07-02","arxiv_id":"2507.01509","n_code_links":1,"syntology":null},{"paper":"/paper/mambattention-mamba-with-multi-head-attention","title":"MambAttention: Mamba with Multi-Head Attention for Generalizable Single-Channel Speech Enhancement","date":"2025-07-01","arxiv_id":"2507.00966","n_code_links":2,"syntology":null},{"paper":"/paper/mamba-fetrack-v2-revisiting-state-space-model","title":"Mamba-FETrack V2: Revisiting State Space Model for Frame-Event based Visual Object Tracking","date":"2025-06-30","arxiv_id":"2506.23783","n_code_links":1,"syntology":null},{"paper":"/paper/eamamba-efficient-all-around-vision-state","title":"EAMamba: Efficient All-Around Vision State Space Model for Image Restoration","date":"2025-06-27","arxiv_id":"2506.22246","n_code_links":1,"syntology":{"ran":4,"of":15,"unverified":11,"pointer_only":15}},{"paper":null,"title":"EAGLE: An Efficient Global Attention Lesion Segmentation Model for Hepatic Echinococcosis","date":"2025-06-25","arxiv_id":"2506.20333","n_code_links":0,"syntology":null},{"paper":null,"title":"FlightKooba: A Fast Interpretable FTP Model","date":"2025-06-24","arxiv_id":"2506.19885","n_code_links":0,"syntology":null},{"paper":null,"title":"JCAPT: A Joint Modeling Approach for CAPT","date":"2025-06-24","arxiv_id":"2506.19315","n_code_links":0,"syntology":null},{"paper":null,"title":"Memba: Membrane-driven Parameter-Efficient Fine-Tuning for Mamba","date":"2025-06-22","arxiv_id":"2506.18184","n_code_links":0,"syntology":null},{"paper":"/paper/vmra-mar-an-asymmetry-aware-temporal","title":"VMRA-MaR: An Asymmetry-Aware Temporal Framework for Longitudinal Breast Cancer Risk Prediction","date":"2025-06-20","arxiv_id":"2506.17412","n_code_links":1,"syntology":null},{"paper":null,"title":"EDNet: A Distortion-Agnostic Speech Enhancement Framework with Gating Mamba Mechanism and Phase Shift-Invariant Training","date":"2025-06-19","arxiv_id":"2506.16231","n_code_links":0,"syntology":null},{"paper":null,"title":"FADPNet: Frequency-Aware Dual-Path Network for Face Super-Resolution","date":"2025-06-17","arxiv_id":"2506.14121","n_code_links":0,"syntology":null},{"paper":null,"title":"MT-PCR: A Hybrid Mamba-Transformer with Spatial Serialization for Hierarchical Point Cloud Registration","date":"2025-06-16","arxiv_id":"2506.13183","n_code_links":0,"syntology":null},{"paper":null,"title":"Scaling Algorithm Distillation for Continuous Control with Mamba","date":"2025-06-16","arxiv_id":"2506.13892","n_code_links":0,"syntology":null},{"paper":null,"title":"Stereo sound event localization and detection based on PSELDnet pretraining and BiMamba sequence modeling","date":"2025-06-16","arxiv_id":"2506.13455","n_code_links":0,"syntology":null},{"paper":"/paper/dart-differentiable-dynamic-adaptive-region","title":"DART: Differentiable Dynamic Adaptive Region Tokenizer for Vision Transformer and Mamba","date":"2025-06-12","arxiv_id":"2506.10390","n_code_links":1,"syntology":null},{"paper":null,"title":"M4V: Multi-Modal Mamba for Text-to-Video Generation","date":"2025-06-12","arxiv_id":"2506.10915","n_code_links":0,"syntology":null},{"paper":null,"title":"Sequential-Parallel Duality in Prefix Scannable Models","date":"2025-06-12","arxiv_id":"2506.10918","n_code_links":0,"syntology":null},{"paper":null,"title":"SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-Shot","date":"2025-06-11","arxiv_id":"2506.09613","n_code_links":0,"syntology":null},{"paper":null,"title":"ECMNet:Lightweight Semantic Segmentation with Efficient CNN-Mamba Network","date":"2025-06-10","arxiv_id":"2506.08629","n_code_links":0,"syntology":null},{"paper":"/paper/inceptionmamba-an-efficient-hybrid-network","title":"InceptionMamba: An Efficient Hybrid Network with Large Band Convolution and Bottleneck Mamba","date":"2025-06-10","arxiv_id":"2506.08735","n_code_links":1,"syntology":null},{"paper":null,"title":"MLVTG: Mamba-Based Feature Alignment and LLM-Driven Purification for Multi-Modal Video Temporal Grounding","date":"2025-06-10","arxiv_id":"2506.08512","n_code_links":0,"syntology":null},{"paper":"/paper/sema-a-scalable-and-efficient-mamba-like","title":"SEMA: a Scalable and Efficient Mamba like Attention via Token Localization and Averaging","date":"2025-06-10","arxiv_id":"2506.08297","n_code_links":0,"syntology":{"ran":4,"of":13,"unverified":9,"pointer_only":13}},{"paper":null,"title":"M2Restore: Mixture-of-Experts-based Mamba-CNN Fusion Framework for All-in-One Image Restoration","date":"2025-06-09","arxiv_id":"2506.07814","n_code_links":0,"syntology":null},{"paper":"/paper/flood-damagesense-multimodal-mamba-with","title":"Flood-DamageSense: Multimodal Mamba with Multitask Learning for Building Flood Damage Assessment using SAR Remote Sensing Imagery","date":"2025-06-07","arxiv_id":"2506.06667","n_code_links":1,"syntology":null},{"paper":null,"title":"DM-SegNet: Dual-Mamba Architecture for 3D Medical Image Segmentation with Global Context Modeling","date":"2025-06-05","arxiv_id":"2506.05297","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/mamba","name":"Mamba","papers":594},{"task":"/task/state-space-models","name":"State Space Models","papers":163},{"task":"/task/computational-efficiency","name":"Computational Efficiency","papers":62},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":57},{"task":"/task/segmentation","name":"Segmentation","papers":37},{"task":"/task/image-classification","name":"Image Classification","papers":32},{"task":"/task/image-segmentation","name":"Image Segmentation","papers":32},{"task":"/task/language-modelling","name":"Language Modelling","papers":32},{"task":"/task/object-detection","name":"Object Detection","papers":32},{"task":"/task/language-modeling","name":"Language Modeling","papers":30},{"task":"/task/object-detection-1","name":"object-detection","papers":30},{"task":"/task/decoder","name":"Decoder","papers":29},{"task":"/task/medical-image-segmentation","name":"Medical Image Segmentation","papers":29},{"task":"/task/image-classification","name":"image-classification","papers":29},{"task":"/task/time-series-1","name":"Time Series","papers":24},{"task":"/task/super-resolution","name":"Super-Resolution","papers":23},{"task":null,"name":"GPU","papers":20},{"task":"/task/autonomous-driving","name":"Autonomous Driving","papers":19},{"task":"/task/time-series-forecasting","name":"Time Series Forecasting","papers":19},{"task":"/task/denoising","name":"Denoising","papers":17}],"tasks_shown":20,"n_tasks":384,"usage_by_year":[{"year":"2023","papers":1},{"year":"2024","papers":307},{"year":"2025","papers":291}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/mamba"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}