Papers › VideoChat: Chat-Centric Video Understanding

VideoChat: Chat-Centric Video Understanding

10 May 2023arXiv:2305.06355archive 2025-07-28

Kunchang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, LiMin Wang, Yu Qiao

In this paper, we initiate an attempt of developing an end-to-end chat-centric video understanding system, coined as VideoChat. It integrates video foundation models and large language models via a learnable neural interface, excelling in spatiotemporal reasoning, event localization, and causal relationship inference. To instructively tune this system, we build a video-centric instruction dataset, composed of thousands of videos associated with detailed descriptions and conversations. This dataset emphasizes spatiotemporal reasoning and captures causal relationships, providing a valuable asset for training our chat-centric video understanding system. Preliminary qualitative experiments demonstrate the potential of our system across a broad spectrum of video applications, which could serve as a simple prototype system for future research on chat-centric video understanding. Access our code and data at https://github.com/OpenGVLab/Ask-Anything

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

opengvlab/ask-anything officialmentioned in papermentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Question AnsweringVideo Question AnsweringVideo UnderstandingVideo-based Generative Performance BenchmarkingVideo-based Generative Performance Benchmarking (Consistency)Video-based Generative Performance Benchmarking (Contextual Understanding)Video-based Generative Performance Benchmarking (Correctness of Information)Video-based Generative Performance Benchmarking (Detail Orientation))Video-based Generative Performance Benchmarking (Temporal Understanding)Zero-Shot Video Question Answer

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Question Answering NExT-QA (Open-ended VideoQA) VideoChat Accuracy 56.6 #3 of 6 Archive leaderboard report
Question Answering NExT-QA (Open-ended VideoQA) VideoChat Confidence Score 3.2 #3 of 6 Archive leaderboard report
Video Question Answering ActivityNet-QA Video Chat Accuracy 26.5 #33 of 36 Archive leaderboard report
Video Question Answering ActivityNet-QA Video Chat Confidence score 2.2 #33 of 36 Archive leaderboard report
Video Question Answering MVBench VideoChat Avg. 35.5 #18 of 22 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct Video Chat Consistency 2.24 #21 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct Video Chat Contextual Understanding 2.53 #21 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct Video Chat Correctness of Information 2.23 #21 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct Video Chat Detail Orientation 2.50 #21 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct Video Chat Temporal Understanding 1.94 #21 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct Video Chat mean 2.29 #21 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking (Consistency) VideoInstruct Video Chat gpt-score 2.24 #15 of 18 Archive leaderboard report
Video-based Generative Performance Benchmarking (Contextual Understanding) VideoInstruct Video Chat gpt-score 2.53 #16 of 18 Archive leaderboard report
Video-based Generative Performance Benchmarking (Correctness of Information) VideoInstruct Video Chat gpt-score 2.32 #15 of 18 Archive leaderboard report
Video-based Generative Performance Benchmarking (Detail Orientation)) VideoInstruct Video Chat gpt-score 2.50 #15 of 18 Archive leaderboard report
Video-based Generative Performance Benchmarking (Temporal Understanding) VideoInstruct Video Chat gpt-score 1.94 #17 of 18 Archive leaderboard report
Zero-Shot Video Question Answer ActivityNet-QA Video Chat Accuracy 26.5 #26 of 28 Archive leaderboard report
Zero-Shot Video Question Answer ActivityNet-QA Video Chat Confidence Score 2.2 #26 of 28 Archive leaderboard report
Zero-Shot Video Question Answer MSRVTT-QA Video Chat-7B Accuracy 45.0 #28 of 30 Archive leaderboard report
Zero-Shot Video Question Answer MSRVTT-QA Video Chat-7B Confidence Score 2.5 #28 of 30 Archive leaderboard report
Zero-Shot Video Question Answer MSVD-QA Video Chat-7B Accuracy 56.3 #25 of 28 Archive leaderboard report
Zero-Shot Video Question Answer MSVD-QA Video Chat-7B Confidence Score 2.8 #25 of 28 Archive leaderboard report
Zero-Shot Video Question Answer TGIF-QA Video Chat-7B Accuracy 34.4 #14 of 14 Archive leaderboard report
Zero-Shot Video Question Answer TGIF-QA Video Chat-7B Confidence Score 2.3 #14 of 14 Archive leaderboard report
Zero-Shot Video Question Answer VNBench VideoChat2 Accuracy 12.4 #7 of 9 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections