Papers › Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

8 Jun 2023arXiv:2306.05424archive 2025-07-28

Muhammad Maaz, Hanoona Rasheed, Salman Khan, Fahad Shahbaz Khan

Conversation agents fueled by Large Language Models (LLMs) are providing a new way to interact with visual data. While there have been initial attempts for image-based conversation models, this work addresses the under-explored field of \emph{video-based conversation} by introducing Video-ChatGPT. It is a multimodal model that merges a video-adapted visual encoder with an LLM. The resulting model is capable of understanding and generating detailed conversations about videos. We introduce a new dataset of 100,000 video-instruction pairs used to train Video-ChatGPT acquired via manual and semi-automated pipeline that is easily scalable and robust to label noise. We also develop a quantitative evaluation framework for video-based dialogue models to objectively analyze the strengths and weaknesses of video-based dialogue models. Code: https://github.com/mbzuai-oryx/Video-ChatGPT.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

mbzuai-oryx/video-chatgpt officialmentioned in papermentioned on GitHubpytorchCC-BY-4.0 report
qiujihao19/artemis mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Question AnsweringVCGBench-DiverseVideo Question AnsweringVideo UnderstandingVideo-based Generative Performance BenchmarkingVideo-based Generative Performance Benchmarking (Consistency)Video-based Generative Performance Benchmarking (Contextual Understanding)Video-based Generative Performance Benchmarking (Correctness of Information)Video-based Generative Performance Benchmarking (Detail Orientation))Video-based Generative Performance Benchmarking (Temporal Understanding)Zero-Shot Video Question Answer

Datasets

Introduced by this paper, per the archive.

VideoInstruct

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Question Answering NExT-QA (Open-ended VideoQA) Video-ChatGPT Accuracy 54.6 #5 of 6 Archive leaderboard report
Question Answering NExT-QA (Open-ended VideoQA) Video-ChatGPT Confidence Score 3.2 #5 of 6 Archive leaderboard report
VCGBench-Diverse VideoInstruct Video-ChatGPT Consistency 2.06 #6 of 6 Archive leaderboard report
VCGBench-Diverse VideoInstruct Video-ChatGPT Contextual Understanding 2.46 #6 of 6 Archive leaderboard report
VCGBench-Diverse VideoInstruct Video-ChatGPT Correctness of Information 2.07 #6 of 6 Archive leaderboard report
VCGBench-Diverse VideoInstruct Video-ChatGPT Dense Captioning 0.89 #6 of 6 Archive leaderboard report
VCGBench-Diverse VideoInstruct Video-ChatGPT Detail Orientation 2.42 #6 of 6 Archive leaderboard report
VCGBench-Diverse VideoInstruct Video-ChatGPT Reasoning 3.60 #6 of 6 Archive leaderboard report
VCGBench-Diverse VideoInstruct Video-ChatGPT Spatial Understanding 2.25 #6 of 6 Archive leaderboard report
VCGBench-Diverse VideoInstruct Video-ChatGPT Temporal Understanding 1.39 #6 of 6 Archive leaderboard report
VCGBench-Diverse VideoInstruct Video-ChatGPT mean 2.08 #6 of 6 Archive leaderboard report
Video Question Answering ActivityNet-QA Video-ChatGPT Accuracy 35.2 #29 of 36 Archive leaderboard report
Video Question Answering ActivityNet-QA Video-ChatGPT Confidence score 2.7 #29 of 36 Archive leaderboard report
Video Question Answering MVBench Video-ChatGPT Avg. 32.7 #20 of 22 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct Video-ChatGPT Consistency 2.37 #20 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct Video-ChatGPT Contextual Understanding 2.62 #20 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct Video-ChatGPT Correctness of Information 2.4 #20 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct Video-ChatGPT Detail Orientation 2.52 #20 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct Video-ChatGPT Temporal Understanding 1.98 #20 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct Video-ChatGPT mean 2.38 #20 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking (Consistency) VideoInstruct Video-ChatGPT gpt-score 2.37 #14 of 18 Archive leaderboard report
Video-based Generative Performance Benchmarking (Contextual Understanding) VideoInstruct Video-ChatGPT gpt-score 2.62 #15 of 18 Archive leaderboard report
Video-based Generative Performance Benchmarking (Correctness of Information) VideoInstruct Video-ChatGPT gpt-score 2.40 #14 of 18 Archive leaderboard report
Video-based Generative Performance Benchmarking (Detail Orientation)) VideoInstruct Video-ChatGPT gpt-score 2.52 #14 of 18 Archive leaderboard report
Video-based Generative Performance Benchmarking (Temporal Understanding) VideoInstruct Video-ChatGPT gpt-score 1.98 #15 of 18 Archive leaderboard report
Zero-Shot Video Question Answer ActivityNet-QA Video-ChatGPT Accuracy 35.2 #24 of 28 Archive leaderboard report
Zero-Shot Video Question Answer ActivityNet-QA Video-ChatGPT Confidence Score 2.7 #24 of 28 Archive leaderboard report
Zero-Shot Video Question Answer MSRVTT-QA Video-ChatGPT-7B Accuracy 49.3 #27 of 30 Archive leaderboard report
Zero-Shot Video Question Answer MSRVTT-QA Video-ChatGPT-7B Confidence Score 2.8 #27 of 30 Archive leaderboard report
Zero-Shot Video Question Answer MSVD-QA Video-ChatGPT-7B Accuracy 64.9 #24 of 28 Archive leaderboard report
Zero-Shot Video Question Answer MSVD-QA Video-ChatGPT-7B Confidence Score 3.3 #24 of 28 Archive leaderboard report
Zero-Shot Video Question Answer TGIF-QA Video-ChatGPT-7B Accuracy 51.4 #12 of 14 Archive leaderboard report
Zero-Shot Video Question Answer TGIF-QA Video-ChatGPT-7B Confidence Score 3.0 #12 of 14 Archive leaderboard report
Zero-Shot Video Question Answer VNBench VideoChatGPT Accuracy 4.1 #9 of 9 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections