{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/video-instruction-tuning-with-synthetic-data","title":"Video Instruction Tuning With Synthetic Data","arxiv_id":"2410.02713","date":"2024-10-03","proceeding":null,"authors":["Yuanhan Zhang","Jinming Wu","Wei Li","Bo Li","Zejun Ma","Ziwei Liu","Chunyuan Li"],"abstract":"The development of video large multimodal models (LMMs) has been hindered by the difficulty of curating large amounts of high-quality raw data from the web. To address this, we propose an alternative approach by creating a high-quality synthetic dataset specifically for video instruction-following, namely LLaVA-Video-178K. This dataset includes key tasks such as detailed captioning, open-ended question-answering (QA), and multiple-choice QA. By training on this dataset, in combination with existing visual instruction tuning data, we introduce LLaVA-Video, a new video LMM. Our experiments demonstrate that LLaVA-Video achieves strong performance across various video benchmarks, highlighting the effectiveness of our dataset. We plan to release the dataset, its generation pipeline, and the model checkpoints.","url_abs":"https://arxiv.org/abs/2410.02713v2","url_pdf":"https://arxiv.org/pdf/2410.02713v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"3d-question-answering-3d-qa","task_name":"3D Question Answering (3D-QA)"},{"task_slug":"instruction-following","task_name":"Instruction Following"},{"task_slug":"multiple-choice","task_name":"Multiple-choice"},{"task_slug":"open-question","task_name":"Open-Ended Question Answering"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"video-question-answering","task_name":"Video Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"},{"task_slug":"zeroshot-video-question-answer","task_name":"Zero-Shot Video Question Answer"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/on-implicitqa","task":"","dataset":"ImplicitQA","model":"LLaVA-Video - 7B","rank_in_archive_order":6,"of":7,"metrics":{"Average Accuracy":"42.1","Macro Average Accuracy":"46.3"},"uses_additional_data":false},{"leaderboard":"/sota/3d-question-answering-3d-qa-on-sqa3d","task":"3D Question Answering (3D-QA)","dataset":"SQA3D","model":"LLaVA-Video","rank_in_archive_order":7,"of":13,"metrics":{"Exact Match":"48.5"},"uses_additional_data":false},{"leaderboard":"/sota/video-question-answering-on-next-qa","task":"Video Question Answering","dataset":"NExT-QA","model":"LLaVA-Video","rank_in_archive_order":7,"of":47,"metrics":{"Accuracy":"83.2"},"uses_additional_data":false},{"leaderboard":"/sota/video-question-answering-on-tvbench","task":"Video Question Answering","dataset":"TVBench","model":"LLaVA-Video 72B","rank_in_archive_order":13,"of":28,"metrics":{"Average Accuracy":"50.0"},"uses_additional_data":false},{"leaderboard":"/sota/video-question-answering-on-tvbench","task":"Video Question Answering","dataset":"TVBench","model":"LLaVA-Video 7B","rank_in_archive_order":17,"of":28,"metrics":{"Average Accuracy":"45.6"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-vqa-on-vlm2-bench","task":"Visual Question Answering (VQA)","dataset":"VLM2-Bench","model":"LLaVA-Video-7B","rank_in_archive_order":6,"of":9,"metrics":{"Average Score on VLM2-bench (9 subtasks)":"43.32","GC-mat":"18.53","GC-trk":"12.79","OC-cnt":"62.47","OC-cpr":"54.72","OC-grp":"28.50","PC-VID":"59.00","PC-cnt":"66.91","PC-cpr":"62.00","PC-grp":"25.00"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-video-question-answer-on-zero-shot","task":"Zero-Shot Video Question Answer","dataset":"Zero-shot Video Question Answering on LongVideoBench","model":"LLaVA-Video","rank_in_archive_order":4,"of":4,"metrics":{"Accuracy (% )":"61.9"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2410.02713","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}