Papers › Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Albert Gu, Tri Dao
Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module. Many subquadratic-time architectures such as linear attention, gated convolution and recurrent models, and structured state space models (SSMs) have been developed to address Transformers' computational inefficiency on long sequences, but they have not performed as well as attention on important modalities such as language. We identify that a key weakness of such models is their inability to perform content-based reasoning, and make several improvements. First, simply letting the SSM parameters be functions of the input addresses their weakness with discrete modalities, allowing the model to selectively propagate or forget information along the sequence length dimension depending on the current token. Second, even though this change prevents the use of efficient convolutions, we design a hardware-aware parallel algorithm in recurrent mode. We integrate these selective SSMs into a simplified end-to-end neural network architecture without attention or even MLP blocks (Mamba). Mamba enjoys fast inference (5× higher throughput than Transformers) and linear scaling in sequence length, and its performance improves on real data up to million-length sequences. As a general sequence model backbone, Mamba achieves state-of-the-art performance across several modalities such as language, audio, and genomics. On language modeling, our Mamba-3B model outperforms Transformers of the same size and matches Transformers twice its size, both in pretraining and downstream evaluation.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
For agents, Syntology's MCP tool lists every function and class Syntology harvested from this paper and whether it ran (how to connect): get_harvested_code_for_paper(arxiv_id="2312.00752")
Code
Syntology Ran 18 of 62 code samples harvested from 14 repositories linked to this paper; 44 have no recorded run. Of those that ran: 4 ran · our draft was wrong; 2 ran · fixture could not drive it; 12 ran with no contract checked.
By repository: community (archive-listed): 62 samples from 14 repositories, 18 ran. The run record, sample by sample. “Ran” means executed on a synthesized input, not that the code is correct or reproduces the paper.
35 repositories listed; official and paper-mentioned ones first.
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
62 samples harvested; 18 ran; 0 honoured the contract we drafted; 44 have no recorded run. Read from Syntology's graph 2026-09-24; that is when this build read the record, not when the samples ran.
Licence: 28 of the 62 samples are pointer only, meaning Syntology does not serve that copy's text. This page shows no code text for any sample; each one links to its file in the repository.
Harvested from 14 repositories linked to this paper, official or community; each sample names its own and says which. “Ran” means the sample executed on a synthesized input. It does not mean the output is correct, and nothing here reproduces the paper's results. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code.
Each sample ends with its code_sha256, Syntology's identity for that exact code. An agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.
Repository labels, per sample. official repository: The archive marks this repository official for the paper. named in the paper: The archive records that the paper mentions this repository; it is not marked official. community (archive-listed): In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper. found in paper text by Syntology: Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted. community: Not in the archive's code links for this paper; a community repository Syntology harvested. Samples from a repository marked official are listed first. Licence labels name the repository's licence as recorded at harvest. “Pointer only” means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence label for the reason. File links open the file on GitHub at the default branch, which may have changed since the harvest.
b2b4897f1f260e1e · report
76fa2650a64a7d15 · report
a99d863241faf517 · report
66eac000db0774c9 · report
2aeed50d3dee6266 · report
a58b40a1cca60e8b · report
ad10c82da33733f2 · report
ba026d72b685bee1 · report
c6eb39247d306bdd · report
75f8a49558b26e64 · report
3fcbb53cc7333590 · report
24385f6e21ea88d7 · report
4a0feeebf8ac43a8 · report
3ccd019afab6e874 · report
55bc02c6846634c4 · report
d5e983b865a3413d · report
8558535d4b8ed416 · report
d81ce48746793832 · report
e62f17401cf26b98 · report
0be091c646cafedd · report
6e0d3f18a6d838b9 · report
01433a18b73a69d8 · report
b8656f4ec2336e2c · report
cba45ceb264445f9 · report
47df96dc59196907 · report
b6e3bc590d74b59c · report
b853ef0ddbe24ac6 · report
58b5b2e41748edc1 · report
b8fab631b07b9ceb · report
cced358b4d4bb38d · report
d5649c5229d3c6f2 · report
db4a88f1431b30bd · report
92d457e0ce6c4038 · report
d58967bba23da4e6 · report
ebeb7b6e20a434cf · report
1cbbea4b9de815f6 · report
1bbd9644f0c6f236 · report
2e5e1b8467df7fb2 · report
8f3c3db8de4abd45 · report
6367281f3513cb50 · report
0e3d9394e01f8e81 · report
ad88c39d5b497053 · report
faaf3d02abfa732c · report
546a5f60b7dbd8fe · report
0a8ec528f858d85d · report
37fea50236b0973b · report
bf2818a19b4d3706 · report
5984ddd05552649a · report
be231a069d3f7c48 · report
19e594df117ee0c3 · report
a6bd393ea522c000 · report
10cfd3bf57dd2545 · report
da2a26ac4e6fac71 · report
1eed9cc6741a3c54 · report
2e14c40a76acb232 · report
8227c505af8b8afd · report
98d508bb5b28dd76 · report
ea69e8168673d498 · report
5d474c2e4ca27948 · report
22fcb344f3b02930 · report
bd2682f3bd8b38fc · report
000e6765f1790cbf · report
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Common Sense Reasoning | ARC (Easy) | Mamba-2.8B (0-shot) | Accuracy | 69.7 | #35 of 47 | Archive leaderboard | report |
| Language Modelling | LAMBADA | Mamba-2.8B | Accuracy | 69.2 | #25 of 37 | Archive leaderboard | report |
| Language Modelling | LAMBADA | Mamba-2.8B | Perplexity | 4.23 | #25 of 37 | Archive leaderboard | report |
| Sentence Completion | HellaSwag | Mamba-2.8B | Accuracy | 66.1 | #58 of 89 | Archive leaderboard | report |
| Sentence Completion | HellaSwag | Mamba-1.4B | Accuracy | 59.1 | #61 of 89 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Introduced by this paper: Mamba
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections