Papers › BAMM: Bidirectional Autoregressive Motion Model
BAMM: Bidirectional Autoregressive Motion Model
Ekkasit Pinyoanuntapong, Muhammad Usama Saleem, Pu Wang, Minwoo Lee, Srijan Das, Chen Chen
Generating human motion from text has been dominated by denoising motion models either through diffusion or generative masking process. However, these models face great limitations in usability by requiring prior knowledge of the motion length. Conversely, autoregressive motion models address this limitation by adaptively predicting motion endpoints, at the cost of degraded generation quality and editing capabilities. To address these challenges, we propose Bidirectional Autoregressive Motion Model (BAMM), a novel text-to-motion generation framework. BAMM consists of two key components: (1) a motion tokenizer that transforms 3D human motion into discrete tokens in latent space, and (2) a masked self-attention transformer that autoregressively predicts randomly masked tokens via a hybrid attention masking strategy. By unifying generative masked modeling and autoregressive modeling, BAMM captures rich and bidirectional dependencies among motion tokens, while learning the probabilistic mapping from textual inputs to motion outputs with dynamically-adjusted motion sequence length. This feature enables BAMM to simultaneously achieving high-quality motion generation with enhanced usability and built-in motion editability. Extensive experiments on HumanML3D and KIT-ML datasets demonstrate that BAMM surpasses current state-of-the-art methods in both qualitative and quantitative measures. Our project page is available at https://exitudio.github.io/BAMM-page
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Motion Synthesis | HumanML3D | BAMM | Diversity | 9.717 | #9 of 37 | Archive leaderboard | report |
| Motion Synthesis | HumanML3D | BAMM | FID | 0.055 | #9 of 37 | Archive leaderboard | report |
| Motion Synthesis | HumanML3D | BAMM | Multimodality | 1.687 | #9 of 37 | Archive leaderboard | report |
| Motion Synthesis | HumanML3D | BAMM | R Precision Top3 | 0.814 | #9 of 37 | Archive leaderboard | report |
| Motion Synthesis | KIT Motion-Language | BAMM | Diversity | 11.008 | #7 of 31 | Archive leaderboard | report |
| Motion Synthesis | KIT Motion-Language | BAMM | FID | 0.183 | #7 of 31 | Archive leaderboard | report |
| Motion Synthesis | KIT Motion-Language | BAMM | Multimodality | 1.609 | #7 of 31 | Archive leaderboard | report |
| Motion Synthesis | KIT Motion-Language | BAMM | R Precision Top3 | 0.788 | #7 of 31 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections