Papers › Mirasol3B: A Multimodal Autoregressive model for time-aligned and contextual modalities

Mirasol3B: A Multimodal Autoregressive model for time-aligned and contextual modalities

9 Nov 2023CVPR 2024 1arXiv:2311.05698archive 2025-07-28

AJ Piergiovanni, Isaac Noble, Dahun Kim, Michael S. Ryoo, Victor Gomes, Anelia Angelova

One of the main challenges of multimodal learning is the need to combine heterogeneous modalities (e.g., video, audio, text). For example, video and audio are obtained at much higher rates than text and are roughly aligned in time. They are often not synchronized with text, which comes as a global context, e.g., a title, or a description. Furthermore, video and audio inputs are of much larger volumes, and grow as the video length increases, which naturally requires more compute dedicated to these modalities and makes modeling of long-range dependencies harder. We here decouple the multimodal modeling, dividing it into separate, focused autoregressive models, processing the inputs according to the characteristics of the modalities. We propose a multimodal model, called Mirasol3B, consisting of an autoregressive component for the time-synchronized modalities (audio and video), and an autoregressive component for the context modalities which are not necessarily aligned in time but are still sequential. To address the long-sequences of the video-audio inputs, we propose to further partition the video and audio sequences in consecutive snippets and autoregressively process their representations. To that end, we propose a Combiner mechanism, which models the audio-video information jointly within a timeframe. The Combiner learns to extract audio and video features from raw spatio-temporal signals, and then learns to fuse these features producing compact but expressive representations per snippet. Our approach achieves the state-of-the-art on well established multimodal benchmarks, outperforming much larger models. It effectively addresses the high computational demand of media inputs by both learning compact representations, controlling the sequence length of the audio-video feature representations, and modeling their dependencies in time.

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action ClassificationAudio ClassificationVideo Question Answering

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Classification Kinetics-Sounds Mirasol3B Top 1 Accuracy 90.1 #3 of 4 Archive leaderboard report
Audio Classification EPIC-SOUNDS Mirasol3B Accuracy 78.2 #1 of 3 Archive leaderboard report
Audio Classification VGGSound Mirasol3B Top 1 Accuracy 69.8 #1 of 23 Archive leaderboard report
Video Question Answering ActivityNet-QA Mirasol3B Accuracy 51.13 #4 of 36 Archive leaderboard report
Video Question Answering MSRVTT-QA Mirasol3B Accuracy 50.42 #1 of 14 Archive leaderboard report
Video Question Answering NExT-QA Mirasol3B Accuracy 72 #30 of 47 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections