Papers › Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image...

Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

14 Nov 2023CVPR 2024 1arXiv:2311.08046archive 2025-07-28

Peng Jin, Ryuichi Takanobu, Wancai Zhang, Xiaochun Cao, Li Yuan

Large language models have demonstrated impressive universal capabilities across a wide range of open-ended tasks and have extended their utility to encompass multimodal conversations. However, existing methods encounter challenges in effectively handling both image and video understanding, particularly with limited visual tokens. In this work, we introduce Chat-UniVi, a Unified Vision-language model capable of comprehending and engaging in conversations involving images and videos through a unified visual representation. Specifically, we employ a set of dynamic visual tokens to uniformly represent images and videos. This representation framework empowers the model to efficiently utilize a limited number of visual tokens to simultaneously capture the spatial details necessary for images and the comprehensive temporal relationship required for videos. Moreover, we leverage a multi-scale representation, enabling the model to perceive both high-level semantic concepts and low-level visual details. Notably, Chat-UniVi is trained on a mixed dataset containing both images and videos, allowing direct application to tasks involving both mediums without requiring any modifications. Extensive experimental results demonstrate that Chat-UniVi consistently outperforms even existing methods exclusively designed for either images or videos. Code is available at https://github.com/PKU-YuanGroup/Chat-UniVi.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

pku-yuangroup/chat-univi officialmentioned in papermentioned on GitHubpytorchApache-2.0 report
pku-yuangroup/video-bench mentioned on GitHub report
skyworkai/moe-plus-plus mentioned on GitHubpytorch report
skyworkai/moh mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Language ModelingLanguage ModellingScience Question AnsweringVCGBench-DiverseVideo Question AnsweringVideo UnderstandingVideo-based Generative Performance BenchmarkingVideo-based Generative Performance Benchmarking (Consistency)Video-based Generative Performance Benchmarking (Contextual Understanding)Video-based Generative Performance Benchmarking (Correctness of Information)Video-based Generative Performance Benchmarking (Correctness of Information) on VideoInstructVideo-based Generative Performance Benchmarking (Detail Orientation))Video-based Generative Performance Benchmarking (Temporal Understanding)Zero-Shot Video Question Answer

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Science Question Answering ScienceQA Chat-UniVi-13B Avg. Accuracy 90.99 #5 of 10 Archive leaderboard report
Science Question Answering ScienceQA Chat-UniVi-13B Grades 1-6 91.19 #5 of 10 Archive leaderboard report
Science Question Answering ScienceQA Chat-UniVi-13B Grades 7-12 90.64 #5 of 10 Archive leaderboard report
Science Question Answering ScienceQA Chat-UniVi-13B Image Context 88.05 #5 of 10 Archive leaderboard report
Science Question Answering ScienceQA Chat-UniVi-13B Language Science 88.91 #5 of 10 Archive leaderboard report
Science Question Answering ScienceQA Chat-UniVi-13B Natural Science 90.41 #5 of 10 Archive leaderboard report
Science Question Answering ScienceQA Chat-UniVi-13B No Context 90.94 #5 of 10 Archive leaderboard report
Science Question Answering ScienceQA Chat-UniVi-13B Social Science 95.05 #5 of 10 Archive leaderboard report
Science Question Answering ScienceQA Chat-UniVi-13B Text Context 89.64 #5 of 10 Archive leaderboard report
VCGBench-Diverse VideoInstruct Chat-UniVi Consistency 2.36 #2 of 6 Archive leaderboard report
VCGBench-Diverse VideoInstruct Chat-UniVi Contextual Understanding 2.66 #2 of 6 Archive leaderboard report
VCGBench-Diverse VideoInstruct Chat-UniVi Correctness of Information 2.29 #2 of 6 Archive leaderboard report
VCGBench-Diverse VideoInstruct Chat-UniVi Dense Captioning 1.33 #2 of 6 Archive leaderboard report
VCGBench-Diverse VideoInstruct Chat-UniVi Detail Orientation 2.56 #2 of 6 Archive leaderboard report
VCGBench-Diverse VideoInstruct Chat-UniVi Reasoning 3.59 #2 of 6 Archive leaderboard report
VCGBench-Diverse VideoInstruct Chat-UniVi Spatial Understanding 2.36 #2 of 6 Archive leaderboard report
VCGBench-Diverse VideoInstruct Chat-UniVi Temporal Understanding 1.56 #2 of 6 Archive leaderboard report
VCGBench-Diverse VideoInstruct Chat-UniVi mean 2.29 #2 of 6 Archive leaderboard report
Video Question Answering ActivityNet-QA Chat-UniVi-13B Accuracy 46.4 #13 of 36 Archive leaderboard report
Video Question Answering ActivityNet-QA Chat-UniVi-13B Confidence score 3.3 #13 of 36 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct Chat-UniVi Consistency 2.81 #14 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct Chat-UniVi Contextual Understanding 3.46 #14 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct Chat-UniVi Correctness of Information 2.89 #14 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct Chat-UniVi Detail Orientation 2.91 #14 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct Chat-UniVi Temporal Understanding 2.39 #14 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking VideoInstruct Chat-UniVi mean 2.99 #14 of 23 Archive leaderboard report
Video-based Generative Performance Benchmarking (Consistency) VideoInstruct Chat-UniVi gpt-score 2.81 #6 of 18 Archive leaderboard report
Video-based Generative Performance Benchmarking (Contextual Understanding) VideoInstruct Chat-UniVi gpt-score 3.46 #10 of 18 Archive leaderboard report
Video-based Generative Performance Benchmarking (Correctness of Information) VideoInstruct Chat-UniVi gpt-score 2.89 #10 of 18 Archive leaderboard report
Video-based Generative Performance Benchmarking (Detail Orientation)) VideoInstruct Chat-UniVi gpt-score 2.91 #10 of 18 Archive leaderboard report
Video-based Generative Performance Benchmarking (Temporal Understanding) VideoInstruct Chat-UniVi gpt-score 2.39 #11 of 18 Archive leaderboard report
Zero-Shot Video Question Answer ActivityNet-QA Chat-UniVi-13B Accuracy 46.4 #17 of 28 Archive leaderboard report
Zero-Shot Video Question Answer ActivityNet-QA Chat-UniVi-13B Confidence Score 3.6 #17 of 28 Archive leaderboard report
Zero-Shot Video Question Answer ActivityNet-QA Chat-UniVi Accuracy 46.1 #19 of 28 Archive leaderboard report
Zero-Shot Video Question Answer ActivityNet-QA Chat-UniVi Confidence Score 3.3 #19 of 28 Archive leaderboard report
Zero-Shot Video Question Answer MSRVTT-QA Chat-UniVi-7B Accuracy 55.0 #22 of 30 Archive leaderboard report
Zero-Shot Video Question Answer MSRVTT-QA Chat-UniVi-7B Confidence Score 3.1 #22 of 30 Archive leaderboard report
Zero-Shot Video Question Answer MSVD-QA Chat-UniVi-7B Accuracy 69.3 #21 of 28 Archive leaderboard report
Zero-Shot Video Question Answer MSVD-QA Chat-UniVi-7B Confidence Score 3.7 #21 of 28 Archive leaderboard report
Zero-Shot Video Question Answer TGIF-QA Chat-UniVi-7B Accuracy 69.0 #10 of 14 Archive leaderboard report
Zero-Shot Video Question Answer TGIF-QA Chat-UniVi-7B Confidence Score 3.8 #10 of 14 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

SET

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections