Papers › Video Instruction Tuning With Synthetic Data

Video Instruction Tuning With Synthetic Data

3 Oct 2024arXiv:2410.02713archive 2025-07-28

Yuanhan Zhang, Jinming Wu, Wei Li, Bo Li, Zejun Ma, Ziwei Liu, Chunyuan Li

The development of video large multimodal models (LMMs) has been hindered by the difficulty of curating large amounts of high-quality raw data from the web. To address this, we propose an alternative approach by creating a high-quality synthetic dataset specifically for video instruction-following, namely LLaVA-Video-178K. This dataset includes key tasks such as detailed captioning, open-ended question-answering (QA), and multiple-choice QA. By training on this dataset, in combination with existing visual instruction tuning data, we introduce LLaVA-Video, a new video LMM. Our experiments demonstrate that LLaVA-Video achieves strong performance across various video benchmarks, highlighting the effectiveness of our dataset. We plan to release the dataset, its generation pipeline, and the model checkpoints.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

3D Question Answering (3D-QA)Instruction FollowingMultiple-choiceOpen-Ended Question AnsweringQuestion AnsweringVideo Question AnsweringVisual Question Answering (VQA)Zero-Shot Video Question Answer

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
ImplicitQA LLaVA-Video - 7B Average Accuracy 42.1 #6 of 7 Archive leaderboard report
ImplicitQA LLaVA-Video - 7B Macro Average Accuracy 46.3 #6 of 7 Archive leaderboard report
3D Question Answering (3D-QA) SQA3D LLaVA-Video Exact Match 48.5 #7 of 13 Archive leaderboard report
Video Question Answering NExT-QA LLaVA-Video Accuracy 83.2 #7 of 47 Archive leaderboard report
Video Question Answering TVBench LLaVA-Video 72B Average Accuracy 50.0 #13 of 28 Archive leaderboard report
Video Question Answering TVBench LLaVA-Video 7B Average Accuracy 45.6 #17 of 28 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench LLaVA-Video-7B Average Score on VLM2-bench (9 subtasks) 43.32 #6 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench LLaVA-Video-7B GC-mat 18.53 #6 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench LLaVA-Video-7B GC-trk 12.79 #6 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench LLaVA-Video-7B OC-cnt 62.47 #6 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench LLaVA-Video-7B OC-cpr 54.72 #6 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench LLaVA-Video-7B OC-grp 28.50 #6 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench LLaVA-Video-7B PC-VID 59.00 #6 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench LLaVA-Video-7B PC-cnt 66.91 #6 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench LLaVA-Video-7B PC-cpr 62.00 #6 of 9 Archive leaderboard report
Visual Question Answering (VQA) VLM2-Bench LLaVA-Video-7B PC-grp 25.00 #6 of 9 Archive leaderboard report
Zero-Shot Video Question Answer Zero-shot Video Question Answering on LongVideoBench LLaVA-Video Accuracy (% ) 61.9 #4 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections