Datasets › IntentQA

IntentQA

Introduced by Jiapeng Li et al. in IntentQA: Context-aware Video Intent Reasoning16 Aug 2023 archive 2025-07-28

We contribute an IntentQA dataset with diverse intents in daily social activities.

We utilize NExT-QA as the source dataset to construct our dataset. NExT-QA dataset is a comprehensive VideoQA dataset with rich natural daily social activities and detailed QA annotations. Originally, the NExT-QA dataset categorizes itself into three types, i.e., Causal, Temporal, Descriptive. We select the inference QA types, i.e., Causal and Temporal, rather than the factoid Descriptive, to build our IntentQA dataset. Particularly, we select both the Causal Why and Causal How subtypes under Causal, and the Temporal Previous and Temporal Next subtypes under Temporal. The Causal Why (CW) QA usually takes the form of ‘Why [action]? For [intent]’, with the key action appearing in the question and the intent in the answer. On the contrary, the Causal How (CH) QA usually takes the form of ‘How [intent]? By [action]’, with the key action appearing in the answer and the intent in the question. The Temporal Previous (TP) QA usually takes the form of ‘What [action A] before [action B]? ’, while the Temporal Next (TN) QA takes the form of ‘What [action B] after [action A]? ’. In the TP&TN QA, the intent is not explicitly expressed in the question nor answer, but is the implicit causal factor linking the two sequential actions.

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Zero-Shot Video Question Answer IntentQA ENTER Accuracy 71.5 ENTER: Event Based Interpretable Reasoning for VideoQA — 13 Compare
Video Question Answering IntentQA VideoChat2_HD_mistral Accuarcy 83.4 MVBench: A Comprehensive Multi-modal Video Understanding... opengvlab/ask-anything +2 6 Compare

Papers archive 2025-07-28

15 shown of 15 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 27. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
ENTER: Event Based Interpretable Reasoning for VideoQA 0 1 24 Jan 2025 not harvested
VidCtx: Context-aware Video Question Answering with Image Models 1 1 23 Dec 2024 not harvested
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models 1 1 17 Nov 2024 ran 5 of 11 samples (6 unverified)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models 1 1 22 Jul 2024 ran 4 of 6 samples (2 unverified; 6 pointer-only for licence)
Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA 1 1 13 Jun 2024 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos 1 1 29 May 2024 ran 11 of 11 samples (0 unverified)
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM 1 1 27 Mar 2024 ran 3 of 3 samples (0 unverified)
Language Repository for Long Video Understanding 1 1 21 Mar 2024 ran 9 of 9 samples (0 unverified)
A Simple LLM Framework for Long-Range Video Question-Answering 1 2 28 Dec 2023 ran 6 of 6 samples (0 unverified)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark 3 2 28 Nov 2023 ran 7 of 10 samples (3 unverified)
Mistral 7B 6 1 10 Oct 2023 ran 9 of 11 samples (2 unverified; 1 pointer-only for licence)
Self-Chained Image-Language Model for Video Localization and Question Answering 1 1 11 May 2023 ran 6 of 7 samples (1 unverified; 7 pointer-only for licence)
IntentQA: Context-aware Video Intent Reasoning 1 2 1 Jan 2023 not harvested
Video Graph Transformer for Video Question Answering 1 1 12 Jul 2022 ran 9 of 14 samples (5 unverified)
Video as Conditional Graph Hierarchy for Multi-Granular Question Answering 1 1 12 Dec 2021 ran 1 of 6 samples (5 unverified)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • IntentQA

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections