Datasets › NExT-QA

NExT-QA

Introduced by Junbin Xiao et al. in NExT-QA: Next Phase of Question-Answering to Explaining Temporal Actions19 Jun 2021 archive 2025-07-28

NExT-QA is a VideoQA benchmark targeting the explanation of video contents. It challenges QA models to reason about the causal and temporal actions and understand the rich object interactions in daily activities, e.g., "why is the boy crying?" and "How does the lady react after the boy fall backward?". It supports both multi-choice and generative open-ended QA tasks. The videos are untrimmed and the questions usually invoke local video contents for answers.

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

30 shown of 64 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 174. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering 1 1 25 Apr 2025 not harvested
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding 1 3 17 Apr 2025 not harvested
Agentic Keyframe Search for Video Question Answering 1 1 20 Mar 2025 ran 3 of 8 samples (5 unverified)
BIMBA: Selective-Scan Compression for Long-Range Video Question Answering 1 1 12 Mar 2025 not harvested
ENTER: Event Based Interpretable Reasoning for VideoQA 0 1 24 Jan 2025 not harvested
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding 1 1 22 Jan 2025 ran 5 of 14 samples (9 unverified)
VidCtx: Context-aware Video Question Answering with Image Models 1 1 23 Dec 2024 not harvested
LinVT: Empower Your Image-level Large Language Model to Understand Videos 1 1 6 Dec 2024 ran 4 of 12 samples (8 unverified; 12 pointer-only for licence)
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling 1 1 6 Dec 2024 ran 1 of 9 samples (8 unverified)
NVILA: Efficient Frontier Visual Language Models 2 1 5 Dec 2024 not harvested
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models 1 1 17 Nov 2024 ran 5 of 11 samples (6 unverified)
Video Instruction Tuning With Synthetic Data 0 1 3 Oct 2024 not harvested
Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution 1 1 19 Sep 2024 ran 3 of 6 samples (3 unverified)
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution 8 1 18 Sep 2024 ran 8 of 12 samples (4 unverified)
LongVILA: Scaling Long-Context Visual Language Models for Long Videos 1 1 19 Aug 2024 not harvested
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 1 1 9 Aug 2024 not harvested
LLaVA-OneVision: Easy Visual Task Transfer 2 2 6 Aug 2024 not harvested
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models 1 1 22 Jul 2024 ran 4 of 6 samples (2 unverified; 6 pointer-only for licence)
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models 3 3 10 Jul 2024 ran 2 of 4 samples (2 unverified; 4 pointer-only for licence)
Tarsier: Recipes for Training and Evaluating Large Video Description Models 1 1 30 Jun 2024 ran 2 of 2 samples (0 unverified)
Long Context Transfer from Language to Vision 2 1 24 Jun 2024 ran 5 of 5 samples (0 unverified; 5 pointer-only for licence)
Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA 1 1 13 Jun 2024 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs 3 1 11 Jun 2024 ran 7 of 17 samples (10 unverified)
DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs 0 1 6 Jun 2024 not harvested
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos 1 1 29 May 2024 ran 11 of 11 samples (0 unverified)
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering 0 1 9 Apr 2024 not harvested
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering 1 1 1 Apr 2024 not harvested
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM 1 2 27 Mar 2024 ran 3 of 3 samples (0 unverified)
Understanding Long Videos with Multimodal Language Models 1 1 25 Mar 2024 ran 4 of 4 samples (0 unverified; 1 pointer-only for licence)
Language Repository for Long Video Understanding 1 1 21 Mar 2024 ran 9 of 9 samples (0 unverified)

The full list of 64 is in the JSON twin.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

MIT

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • NExT-QA
  • NExT-GQA
  • NExT-QA (Open-ended VideoQA)

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections