Papers › MuLTI: Efficient Video-and-Language Understanding with Text-Guided MultiWay-Sampler...
MuLTI: Efficient Video-and-Language Understanding with Text-Guided MultiWay-Sampler and Multiple Choice Modeling
Jiaqi Xu, Bo Liu, Yunkuo Chen, Mengli Cheng, Xing Shi
Video-and-language understanding has a variety of applications in the industry, such as video question answering, text-video retrieval, and multi-label classification. Existing video-and-language understanding methods generally adopt heavy multi-modal encoders and feature fusion modules, which consume high computational costs. Specially, they have difficulty dealing with dense video frames or long text prevalent in industrial applications. This paper proposes MuLTI, a highly accurate and efficient video-and-language understanding model that achieves efficient and effective feature fusion and rapid adaptation to downstream tasks. Specifically, we design a Text-Guided MultiWay-Sampler based on adapt-pooling residual mapping and self-attention modules to sample long sequences and fuse multi-modal features, which reduces the computational costs and addresses performance degradation caused by previous samplers. Therefore, MuLTI can handle longer sequences with limited computational costs. Then, to further enhance the model's performance and fill in the lack of pretraining tasks in the video question answering, we propose a new pretraining task named Multiple Choice Modeling. This task bridges the gap between pretraining and downstream tasks and improves the model's ability to align video and text features. Benefiting from the efficient feature fusion module and the new pretraining task, MuLTI achieves state-of-the-art performance on multiple datasets. Implementation and pretrained models will be released.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Video Retrieval | DiDeMo | MuLTI | text-to-video R@1 | 56.5 | #14 of 40 | Archive leaderboard | report |
| Video Retrieval | DiDeMo | MuLTI | text-to-video R@10 | 87.0 | #14 of 40 | Archive leaderboard | report |
| Video Retrieval | DiDeMo | MuLTI | text-to-video R@5 | 80.2 | #14 of 40 | Archive leaderboard | report |
| Video Retrieval | MSR-VTT-1kA | MuLTI | text-to-video R@1 | 54.7 | #6 of 63 | Archive leaderboard | report |
| Video Retrieval | MSR-VTT-1kA | MuLTI | text-to-video R@10 | 86.0 | #6 of 63 | Archive leaderboard | report |
| Video Retrieval | MSR-VTT-1kA | MuLTI | text-to-video R@5 | 77.7 | #6 of 63 | Archive leaderboard | report |
| Visual Question Answering (VQA) | MSRVTT-QA | MuLTI | Accuracy | 0.478 | #4 of 34 | Archive leaderboard | report |
| Visual Question Answering (VQA) | MSVD-QA | MuLTI | Accuracy | 0.547 | #16 of 36 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections