Datasets › VideoInstruct

VideoInstruct (Video Instruction Dataset)

Introduced by Muhammad Maaz et al. in Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models8 Jun 2023 archive 2025-07-28

Video Instruction Dataset is used to train Video-ChatGPT. It consists of 100,000 high-quality video instruction pairs. employs a combination of human-assisted and semi-automatic annotation techniques, aiming to produce high-quality video instruction data. These methods create question-answer pairs related to

  1. Video summarization
  2. Description-based question-answers (exploring spatial, temporal, relationships, and reasoning concepts)
  3. Creative/generative question-answers

Benchmarks archive 2025-07-28

All 7 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

21 shown of 21 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 30. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models 1 6 17 Nov 2024 ran 5 of 11 samples (6 unverified)
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance 1 7 4 Nov 2024 ran 2 of 9 samples (7 unverified; 3 pointer-only for licence)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models 1 6 22 Jul 2024 ran 4 of 6 samples (2 unverified; 6 pointer-only for licence)
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding 1 7 13 Jun 2024 ran 6 of 8 samples (2 unverified; 8 pointer-only for licence)
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning 1 6 25 Apr 2024 ran 0 of 2 samples (2 unverified; 2 pointer-only for licence)
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens 2 5 4 Apr 2024 ran 2 of 2 samples (0 unverified)
ST-LLM: Large Language Models Are Effective Temporal Learners 1 6 30 Mar 2024 ran 7 of 11 samples (4 unverified; 3 pointer-only for licence)
LITA: Language Instructed Temporal-Localization Assistant 1 1 27 Mar 2024 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM 1 1 27 Mar 2024 ran 3 of 3 samples (0 unverified)
CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios 1 1 7 Mar 2024 ran 6 of 12 samples (6 unverified)
Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback 1 1 6 Feb 2024 not harvested
VTimeLLM: Empower LLM to Grasp Video Moments 1 7 30 Nov 2023 ran 5 of 11 samples (6 unverified; 11 pointer-only for licence)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark 3 13 28 Nov 2023 ran 7 of 10 samples (3 unverified)
LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models 2 2 28 Nov 2023 ran 3 of 4 samples (1 unverified)
Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding 4 7 14 Nov 2023 not harvested
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning 1 13 27 Sep 2023 ran 1 of 4 samples (3 unverified)
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding 1 5 31 Jul 2023 not harvested
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models 2 7 8 Jun 2023 not harvested
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding 4 6 5 Jun 2023 ran 18 of 25 samples (7 unverified; 9 pointer-only for licence)
VideoChat: Chat-Centric Video Understanding 1 6 10 May 2023 not harvested
LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model 3 6 28 Apr 2023 ran 0 of 1 samples (1 unverified; 1 pointer-only for licence)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Creative Commons Attribution 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • VideoInstruct

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections