Datasets › ImplicitQA
ImplicitQA
The ImplicitQA dataset was introduced in the paper ImplicitQA: Going beyond frames towards Implicit Video Reasoning.
Project page: https://swetha5.github.io/ImplicitQA/
ImplicitQA is a novel benchmark specifically designed to test models on implicit reasoning in Video Question Answering (VideoQA). Unlike existing VideoQA benchmarks that primarily focus on questions answerable through explicit visual content (actions, objects, events directly observable within individual frames or short clips), ImplicitQA addresses the need for models to infer motives, causality, and relationships across discontinuous frames. This mirrors human-like understanding of creative and cinematic videos, which often employ storytelling techniques that deliberately omit certain depictions.
The dataset comprises 1,000 meticulously annotated QA pairs derived from over 320 high-quality creative video clips. These QA pairs are systematically categorized into key reasoning dimensions, including:
- Lateral spatial reasoning
- Vertical spatial reasoning
- Relative Depth and proximity
- Viewpoint and visibility
- Motion and trajectory Dynamics
- Causal and motivational reasoning
- Social interactions and Relationships
- Physical and Environmental context
- Inferred counting
The annotations are deliberately challenging, crafted to ensure high quality and to highlight the difficulty of implicit reasoning for current VideoQA models.
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| ImplicitQA | GPT O3 Average Accuracy 64.1 | ImplicitQA: Going beyond frames towards Implicit Video Reasoning | UCF-CRCV/ImplicitQA | 7 | Compare |
Papers archive 2025-07-28
5 shown of 5 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 5. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| ImplicitQA: Going beyond frames towards Implicit Video Reasoning | 1 | 1 | 26 Jun 2025 | not harvested |
| Qwen2.5-VL Technical Report | 4 | 1 | 19 Feb 2025 | ran 2 of 3 samples (1 unverified) |
| Video Instruction Tuning With Synthetic Data | 0 | 1 | 3 Oct 2024 | not harvested |
| Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution | 8 | 1 | 18 Sep 2024 | ran 8 of 12 samples (4 unverified) |
| LLaVA-OneVision: Easy Visual Task Transfer | 2 | 1 | 6 Aug 2024 | not harvested |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- ImplicitQA
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections