{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mamba-nd-selective-state-space-modeling-for","title":"Mamba-ND: Selective State Space Modeling for Multi-Dimensional Data","arxiv_id":"2402.05892","date":"2024-02-08","proceeding":null,"authors":["Shufan Li","Harkanwar Singh","Aditya Grover"],"abstract":"In recent years, Transformers have become the de-facto architecture for sequence modeling on text and a variety of multi-dimensional data, such as images and video. However, the use of self-attention layers in a Transformer incurs prohibitive compute and memory complexity that scales quadratically w.r.t. the sequence length. A recent architecture, Mamba, based on state space models has been shown to achieve comparable performance for modeling text sequences, while scaling linearly with the sequence length. In this work, we present Mamba-ND, a generalized design extending the Mamba architecture to arbitrary multi-dimensional data. Our design alternatively unravels the input data across different dimensions following row-major orderings. We provide a systematic comparison of Mamba-ND with several other alternatives, based on prior multi-dimensional extensions such as Bi-directional LSTMs and S4ND. Empirically, we show that Mamba-ND demonstrates performance competitive with the state-of-the-art on a variety of multi-dimensional benchmarks, including ImageNet-1K classification, HMDB-51 action recognition, and ERA5 weather forecasting.","url_abs":"https://arxiv.org/abs/2402.05892v5","url_pdf":"https://arxiv.org/pdf/2402.05892v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mamba-nd-selective-state-space-modeling-for","repo_url":"https://github.com/jacklishufan/mamba-nd","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"mamba","task_name":"Mamba"},{"task_slug":"state-space-models","task_name":"State Space Models"},{"task_slug":"weather-forecasting","task_name":"Weather Forecasting"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2402.05892","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2402.05892"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jacklishufan/mamba-nd","reach":{"status":"ok"}},{"provenance":"deterministic:regex_extraction","url":"https://github.com/jacklishufan/Mamba-ND","reach":{"status":"ok"}}],"summary":{"ran":7,"unverified":4},"by_repo_kind":{"official":{"samples":11,"ran":7,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":11,"samples":[{"code_sha256_prefix":"bc39b8ed5fefa7eb","entry":"crop_clip","repo":"jacklishufan/Mamba-ND","repo_kind":"official","path":"video_pretraining/functional.py","file_url":"https://github.com/jacklishufan/Mamba-ND/blob/HEAD/video_pretraining/functional.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"bc39b8ed5fefa7eb"}},{"code_sha256_prefix":"7b0ecad70df7f175","entry":"dice","repo":"jacklishufan/Mamba-ND","repo_kind":"official","path":"btcv/trainer.py","file_url":"https://github.com/jacklishufan/Mamba-ND/blob/HEAD/btcv/trainer.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7b0ecad70df7f175"}},{"code_sha256_prefix":"6aef977b44c0feaf","entry":"get_loss_scale_for_deepspeed","repo":"jacklishufan/Mamba-ND","repo_kind":"official","path":"video_pretraining/engines/engine_for_finetuning.py","file_url":"https://github.com/jacklishufan/Mamba-ND/blob/HEAD/video_pretraining/engines/engine_for_finetuning.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"6aef977b44c0feaf"}},{"code_sha256_prefix":"2b53e457d5632b4e","entry":"get_resize_sizes","repo":"jacklishufan/Mamba-ND","repo_kind":"official","path":"video_pretraining/functional.py","file_url":"https://github.com/jacklishufan/Mamba-ND/blob/HEAD/video_pretraining/functional.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"2b53e457d5632b4e"}},{"code_sha256_prefix":"981936c73d3cdd12","entry":"resize_clip","repo":"jacklishufan/Mamba-ND","repo_kind":"official","path":"video_pretraining/functional.py","file_url":"https://github.com/jacklishufan/Mamba-ND/blob/HEAD/video_pretraining/functional.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"981936c73d3cdd12"}},{"code_sha256_prefix":"f188ca6f84532f7f","entry":"spatial_sampling","repo":"jacklishufan/Mamba-ND","repo_kind":"official","path":"video_pretraining/datasets/kinetics.py","file_url":"https://github.com/jacklishufan/Mamba-ND/blob/HEAD/video_pretraining/datasets/kinetics.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f188ca6f84532f7f"}},{"code_sha256_prefix":"fb98703650f02bf6","entry":"train_class_batch","repo":"jacklishufan/Mamba-ND","repo_kind":"official","path":"video_pretraining/engines/engine_for_finetuning_regression.py","file_url":"https://github.com/jacklishufan/Mamba-ND/blob/HEAD/video_pretraining/engines/engine_for_finetuning_regression.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"fb98703650f02bf6"}},{"code_sha256_prefix":"43593092d4dd72a6","entry":"causal_conv1d_fn_ref","repo":"jacklishufan/Mamba-ND","repo_kind":"official","path":"image_classification/src/mamba.py","file_url":"https://github.com/jacklishufan/Mamba-ND/blob/HEAD/image_classification/src/mamba.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"43593092d4dd72a6"}},{"code_sha256_prefix":"66146c617de276d3","entry":"resize_pos_embed","repo":"jacklishufan/Mamba-ND","repo_kind":"official","path":"btcv/networks/mamba.py","file_url":"https://github.com/jacklishufan/Mamba-ND/blob/HEAD/btcv/networks/mamba.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"66146c617de276d3"}},{"code_sha256_prefix":"14b3b13b2e7ee844","entry":"tensor_normalize","repo":"jacklishufan/Mamba-ND","repo_kind":"official","path":"video_pretraining/datasets/kinetics.py","file_url":"https://github.com/jacklishufan/Mamba-ND/blob/HEAD/video_pretraining/datasets/kinetics.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"14b3b13b2e7ee844"}},{"code_sha256_prefix":"c7976ea27edc377a","entry":"train_class_batch","repo":"jacklishufan/Mamba-ND","repo_kind":"official","path":"video_pretraining/engines/engine_for_finetuning.py","file_url":"https://github.com/jacklishufan/Mamba-ND/blob/HEAD/video_pretraining/engines/engine_for_finetuning.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c7976ea27edc377a"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}