Papers › mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

9 Aug 2024arXiv:2408.04840archive 2025-07-28

Jiabo Ye, Haiyang Xu, Haowei Liu, Anwen Hu, Ming Yan, Qi Qian, Ji Zhang, Fei Huang, Jingren Zhou

Multi-modal Large Language Models (MLLMs) have demonstrated remarkable capabilities in executing instructions for a variety of single-image tasks. Despite this progress, significant challenges remain in modeling long image sequences. In this work, we introduce the versatile multi-modal large language model, mPLUG-Owl3, which enhances the capability for long image-sequence understanding in scenarios that incorporate retrieved image-text knowledge, interleaved image-text, and lengthy videos. Specifically, we propose novel hyper attention blocks to efficiently integrate vision and language into a common language-guided semantic space, thereby facilitating the processing of extended multi-image scenarios. Extensive experimental results suggest that mPLUG-Owl3 achieves state-of-the-art performance among models with a similar size on single-image, multi-image, and video benchmarks. Moreover, we propose a challenging long visual sequence evaluation named Distractor Resistance to assess the ability of models to maintain focus amidst distractions. Finally, with the proposed architecture, mPLUG-Owl3 demonstrates outstanding performance on ultra-long visual sequence inputs. We hope that mPLUG-Owl3 can contribute to the development of more efficient and powerful multimodal large language models.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

x-plug/mplug-owl officialmentioned in paperpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Language ModelingLanguage ModellingLarge Language ModelVideo Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Question Answering MVBench mPLUG-Owl3(7B) Avg. 59.5 #8 of 22 Archive leaderboard report
Video Question Answering NExT-QA mPLUG-Owl3(8B) Accuracy 78.6 #18 of 47 Archive leaderboard report
Video Question Answering TVBench mPLUG-Owl3 Average Accuracy 42.2 #21 of 28 Archive leaderboard report
Visual Question Answering MM-Vet mPLUG-Owl3 GPT-4 score 40.1 #111 of 231 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench mPLUG-Owl3-7B Average Score on VLM2-bench (9 subtasks) 37.85 #7 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench mPLUG-Owl3-7B GC-mat 17.37 #7 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench mPLUG-Owl3-7B GC-trk 18.26 #7 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench mPLUG-Owl3-7B OC-cnt 62.97 #7 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench mPLUG-Owl3-7B OC-cpr 49.17 #7 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench mPLUG-Owl3-7B OC-grp 31.00 #7 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench mPLUG-Owl3-7B PC-VID 13.50 #7 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench mPLUG-Owl3-7B PC-cnt 58.86 #7 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench mPLUG-Owl3-7B PC-cpr 63.50 #7 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench mPLUG-Owl3-7B PC-grp 26.00 #7 of 9 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AttentionFocusSoftmax

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections