Papers › BIMBA: Selective-Scan Compression for Long-Range Video Question Answering

BIMBA: Selective-Scan Compression for Long-Range Video Question Answering

12 Mar 2025CVPR 2025 1arXiv:2503.09590archive 2025-07-28

Md Mohaiminul Islam, Tushar Nagarajan, Huiyu Wang, Gedas Bertasius, Lorenzo Torresani

Video Question Answering (VQA) in long videos poses the key challenge of extracting relevant information and modeling long-range dependencies from many redundant frames. The self-attention mechanism provides a general solution for sequence modeling, but it has a prohibitive cost when applied to a massive number of spatiotemporal tokens in long videos. Most prior methods rely on compression strategies to lower the computational cost, such as reducing the input length via sparse frame sampling or compressing the output sequence passed to the large language model (LLM) via space-time pooling. However, these naive approaches over-represent redundant information and often miss salient events or fast-occurring space-time patterns. In this work, we introduce BIMBA, an efficient state-space model to handle long-form videos. Our model leverages the selective scan algorithm to learn to effectively select critical information from high-dimensional video and transform it into a reduced token sequence for efficient LLM processing. Extensive experiments demonstrate that BIMBA achieves state-of-the-art accuracy on multiple long-form VQA benchmarks, including PerceptionTest, NExT-QA, EgoSchema, VNBench, LongVideoBench, and Video-MME. Code, and models are publicly available at https://sites.google.com/view/bimba-mllm.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

md-mohaiminul/BIMBA officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Video Question AnsweringZero-Shot Video Question Answer

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Question Answering NExT-QA BIMBA-LLaVA-Qwen2-7B Accuracy 83.73 #5 of 47 Archive leaderboard report
Video Question Answering Perception Test BIMBA-LLaVA-Qwen2-7B Accuracy (Top-1) 68.51 #2 of 6 Archive leaderboard report
Zero-Shot Video Question Answer EgoSchema (fullset) BIMBA-LLaVA-Qwen2-7B Accuracy 71.14 #1 of 29 Archive leaderboard report
Zero-Shot Video Question Answer VNBench BIMBA-LLaVA-Qwen2-7B Accuracy 77.88 #1 of 9 Archive leaderboard report
Zero-Shot Video Question Answer Video-MME BIMBA-LLaVA-Qwen2-7B Accuracy (%) 64.67 #6 of 11 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections