Methods › Mamba
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Mamba
Introduced by Albert Gu et al. in Mamba: Linear-Time Sequence Modeling with Selective State Spaces
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module. Many subquadratic-time architectures such as linear attention, gated convolution and recurrent models, and structured state space models (SSMs) have been developed to address Transformers’ computational inefficiency on long sequences, but they have not performed as well as attention on important modalities such as language. We identify that a key weakness of such models is their inability to perform content-based reasoning, and make several improvements. First, simply letting the SSM parameters be functions of the input addresses their weakness with discrete modalities, allowing the model to selectively propagate or forget information along the sequence length dimension depending on the current token. Second, even though this change prevents the use of efficient convolutions, we design a hardware-aware parallel algorithm in recurrent mode. We integrate these selective SSMs into a simplified end-to-end neural network architecture without attention or even MLP blocks (Mamba). Mamba enjoys fast inference (5× higher throughput than Transformers) and linear scaling in sequence length, and its performance improves on real data up to million-length sequences. As a general sequence model backbone, Mamba achieves state-of-the-art performance across several modalities such as language, audio, and genomics. On language modeling, our Mamba-3B model outperforms Transformers of the same size and matches Transformers twice its size, both in pre-training and downstream evaluation.
Papers archive 2025-07-28
30 shown of 599, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Differential Mamba 8 Jul 2025 · 1 repository · arXiv:2507.06204Syntology ran 2 of 2 samples · 0 unverified
-
LangMamba: A Language-driven Mamba Framework for Low-dose CT Denoising with Vision-language Models 8 Jul 2025 · 1 repository · arXiv:2507.06140
-
FindRec: Stein-Guided Entropic Flow for Multi-Modal Sequential Recommendation 7 Jul 2025 · 1 repository · arXiv:2507.04651
-
MambaFusion: Height-Fidelity Dense Global Fusion for Multi-modal 3D Object Detection 6 Jul 2025 · 1 repository · arXiv:2507.04369Syntology ran 0 of 10 samples · 10 unverified
-
MVNet: Hyperspectral Remote Sensing Image Classification Based on Hybrid Mamba-Transformer Vision Backbone Architecture 6 Jul 2025 · 1 repository · arXiv:2507.04409
-
Mamba Guided Boundary Prior Matters: A New Perspective for Generalized Polyp Segmentation 2 Jul 2025 · 1 repository · arXiv:2507.01509
-
MambAttention: Mamba with Multi-Head Attention for Generalizable Single-Channel Speech Enhancement 1 Jul 2025 · 2 repositories · arXiv:2507.00966
-
Mamba-FETrack V2: Revisiting State Space Model for Frame-Event based Visual Object Tracking 30 Jun 2025 · 1 repository · arXiv:2506.23783
-
EAMamba: Efficient All-Around Vision State Space Model for Image Restoration 27 Jun 2025 · 1 repository · arXiv:2506.22246Syntology ran 4 of 15 samples · 11 unverified · 15 pointer-only (licence)
-
EAGLE: An Efficient Global Attention Lesion Segmentation Model for Hepatic Echinococcosis 25 Jun 2025 · 0 repositories · arXiv:2506.20333
-
FlightKooba: A Fast Interpretable FTP Model 24 Jun 2025 · 0 repositories · arXiv:2506.19885
-
JCAPT: A Joint Modeling Approach for CAPT 24 Jun 2025 · 0 repositories · arXiv:2506.19315
-
Memba: Membrane-driven Parameter-Efficient Fine-Tuning for Mamba 22 Jun 2025 · 0 repositories · arXiv:2506.18184
-
VMRA-MaR: An Asymmetry-Aware Temporal Framework for Longitudinal Breast Cancer Risk Prediction 20 Jun 2025 · 1 repository · arXiv:2506.17412
-
EDNet: A Distortion-Agnostic Speech Enhancement Framework with Gating Mamba Mechanism and Phase Shift-Invariant Training 19 Jun 2025 · 0 repositories · arXiv:2506.16231
-
FADPNet: Frequency-Aware Dual-Path Network for Face Super-Resolution 17 Jun 2025 · 0 repositories · arXiv:2506.14121
-
MT-PCR: A Hybrid Mamba-Transformer with Spatial Serialization for Hierarchical Point Cloud Registration 16 Jun 2025 · 0 repositories · arXiv:2506.13183
-
Scaling Algorithm Distillation for Continuous Control with Mamba 16 Jun 2025 · 0 repositories · arXiv:2506.13892
-
Stereo sound event localization and detection based on PSELDnet pretraining and BiMamba sequence modeling 16 Jun 2025 · 0 repositories · arXiv:2506.13455
-
DART: Differentiable Dynamic Adaptive Region Tokenizer for Vision Transformer and Mamba 12 Jun 2025 · 1 repository · arXiv:2506.10390
-
M4V: Multi-Modal Mamba for Text-to-Video Generation 12 Jun 2025 · 0 repositories · arXiv:2506.10915
-
Sequential-Parallel Duality in Prefix Scannable Models 12 Jun 2025 · 0 repositories · arXiv:2506.10918
-
SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-Shot 11 Jun 2025 · 0 repositories · arXiv:2506.09613
-
ECMNet:Lightweight Semantic Segmentation with Efficient CNN-Mamba Network 10 Jun 2025 · 0 repositories · arXiv:2506.08629
-
InceptionMamba: An Efficient Hybrid Network with Large Band Convolution and Bottleneck Mamba 10 Jun 2025 · 1 repository · arXiv:2506.08735
-
MLVTG: Mamba-Based Feature Alignment and LLM-Driven Purification for Multi-Modal Video Temporal Grounding 10 Jun 2025 · 0 repositories · arXiv:2506.08512
-
SEMA: a Scalable and Efficient Mamba like Attention via Token Localization and Averaging 10 Jun 2025 · 0 repositories · arXiv:2506.08297Syntology ran 4 of 13 samples · 9 unverified · 13 pointer-only (licence)
-
M2Restore: Mixture-of-Experts-based Mamba-CNN Fusion Framework for All-in-One Image Restoration 9 Jun 2025 · 0 repositories · arXiv:2506.07814
-
Flood-DamageSense: Multimodal Mamba with Multitask Learning for Building Flood Damage Assessment using SAR Remote Sensing Imagery 7 Jun 2025 · 1 repository · arXiv:2506.06667
-
DM-SegNet: Dual-Mamba Architecture for 3D Medical Image Segmentation with Global Context Modeling 5 Jun 2025 · 0 repositories · arXiv:2506.05297
Tasks archive 2025-07-28
20 shown of 384 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
The archive places this method in no collection.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections