Papers › LLaVA-OneVision: Easy Visual Task Transfer

LLaVA-OneVision: Easy Visual Task Transfer

6 Aug 2024arXiv:2408.03326archive 2025-07-28

Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, Chunyuan Li

We present LLaVA-OneVision, a family of open large multimodal models (LMMs) developed by consolidating our insights into data, models, and visual representations in the LLaVA-NeXT blog series. Our experimental results demonstrate that LLaVA-OneVision is the first single model that can simultaneously push the performance boundaries of open LMMs in three important computer vision scenarios: single-image, multi-image, and video scenarios. Importantly, the design of LLaVA-OneVision allows strong transfer learning across different modalities/scenarios, yielding new emerging capabilities. In particular, strong video understanding and cross-scenario capabilities are demonstrated through task transfer from images to videos.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

evolvinglmms-lab/lmms-eval officialmentioned in paperpytorchNOASSERTION report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

3D Question Answering (3D-QA)Multiple-choiceTemporal Relation ExtractionTransfer LearningVideo Question AnsweringVideo UnderstandingVisual Question AnsweringVisual Question Answering (VQA)Zero-Shot Video Question Answer

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
ImplicitQA LLaVA-OneVision - 7B Average Accuracy 43.4 #4 of 7 Archive leaderboard report
ImplicitQA LLaVA-OneVision - 7B Macro Average Accuracy 46.4 #4 of 7 Archive leaderboard report
3D Question Answering (3D-QA) SQA3D LLaVA-NeXT-Video Exact Match 34.2 #13 of 13 Archive leaderboard report
3D Question Answering (3D-QA) ScanQA Test w/ objects LLaVA-NeXT-Video BLEU-4 9.8 #16 of 18 Archive leaderboard report
3D Question Answering (3D-QA) ScanQA Test w/ objects LLaVA-NeXT-Video CIDEr 46.2 #16 of 18 Archive leaderboard report
3D Question Answering (3D-QA) ScanQA Test w/ objects LLaVA-NeXT-Video Exact Match 18.7 #16 of 18 Archive leaderboard report
3D Question Answering (3D-QA) ScanQA Test w/ objects LLaVA-NeXT-Video METEOR 9.1 #16 of 18 Archive leaderboard report
3D Question Answering (3D-QA) ScanQA Test w/ objects LLaVA-NeXT-Video ROUGE 27.8 #16 of 18 Archive leaderboard report
Temporal Relation Extraction Vinoground LLaVA-OneVision-Qwen2-72B Group Score 21.8 #4 of 24 Archive leaderboard report
Temporal Relation Extraction Vinoground LLaVA-OneVision-Qwen2-72B Text Score 48.4 #4 of 24 Archive leaderboard report
Temporal Relation Extraction Vinoground LLaVA-OneVision-Qwen2-72B Video Score 35.2 #4 of 24 Archive leaderboard report
Temporal Relation Extraction Vinoground LLaVA-OneVision-Qwen2-7B Group Score 14.6 #5 of 24 Archive leaderboard report
Temporal Relation Extraction Vinoground LLaVA-OneVision-Qwen2-7B Text Score 41.6 #5 of 24 Archive leaderboard report
Temporal Relation Extraction Vinoground LLaVA-OneVision-Qwen2-7B Video Score 29.4 #5 of 24 Archive leaderboard report
Video Question Answering NExT-QA LLaVA-OV(72B) Accuracy 80.2 #13 of 47 Archive leaderboard report
Video Question Answering NExT-QA LLaVA-OV(7B) Accuracy 79.4 #15 of 47 Archive leaderboard report
Video Question Answering OVBench LLaVA-OneVision (7B) AVG 49.5 #5 of 16 Archive leaderboard report
Visual Question Answering MM-Vet LLaVA-OneVision-72B GPT-4 score 63.7 #26 of 231 Archive leaderboard report
Visual Question Answering MM-Vet LLaVA-OneVision-7B GPT-4 score 57.5 #43 of 231 Archive leaderboard report
Visual Question Answering MM-Vet LLaVA-OneVision-0.5B GPT-4 score 29.1 #212 of 231 Archive leaderboard report
Visual Question Answering V*bench LLaVA-OneVision7B Accuracy 74.46 #5 of 5 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench LLaVA-OneVision-7B Average Score on VLM2-bench (9 subtasks) 39.35 #8 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench LLaVA-OneVision-7B GC-mat 16.60 #8 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench LLaVA-OneVision-7B GC-trk 13.70 #8 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench LLaVA-OneVision-7B OC-cnt 56.17 #8 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench LLaVA-OneVision-7B OC-cpr 47.22 #8 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench LLaVA-OneVision-7B OC-grp 27.50 #8 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench LLaVA-OneVision-7B PC-VID 47.25 #8 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench LLaVA-OneVision-7B PC-cnt 46.67 #8 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench LLaVA-OneVision-7B PC-cpr 62.00 #8 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench LLaVA-OneVision-7B PC-grp 37.00 #8 of 9 Archive leaderboard report
Zero-Shot Video Question Answer VNBench LLaVA-OneVision-72B Accuracy 58.7 #3 of 9 Archive leaderboard report
Zero-Shot Video Question Answer VNBench LLaVA-OneVision-7B Accuracy 51.8 #4 of 9 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections