Datasets › TGIF-QA

TGIF-QA

Introduced by Yunseok Jang et al. in TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering14 Apr 2017 archive 2025-07-28

The TGIF-QA dataset contains 165K QA pairs for the animated GIFs from the TGIF dataset [Li et al. CVPR 2016]. The question & answer pairs are collected via crowdsourcing with a carefully designed user interface to ensure quality. The dataset can be used to evaluate video-based Visual Question Answering techniques.

Source: GitHub Image Source: https://github.com/YunseokJANG/tgif-qa

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

17 shown of 17 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 92. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
LinVT: Empower Your Image-level Large Language Model to Understand Videos 1 1 6 Dec 2024 ran 4 of 12 samples (8 unverified; 12 pointer-only for licence)
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models 1 1 17 Nov 2024 ran 5 of 11 samples (6 unverified)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models 1 1 22 Jul 2024 ran 4 of 6 samples (2 unverified; 6 pointer-only for licence)
Tarsier: Recipes for Training and Evaluating Large Video Description Models 1 1 30 Jun 2024 ran 2 of 2 samples (0 unverified)
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding 1 1 13 Jun 2024 ran 6 of 8 samples (2 unverified; 8 pointer-only for licence)
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning 1 1 25 Apr 2024 ran 0 of 2 samples (2 unverified; 2 pointer-only for licence)
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs 1 1 11 Apr 2024 not harvested
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens 2 1 4 Apr 2024 ran 2 of 2 samples (0 unverified)
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM 1 1 27 Mar 2024 ran 3 of 3 samples (0 unverified)
Elysium: Exploring Object-level Perception in Videos via MLLM 1 1 25 Mar 2024 ran 7 of 8 samples (1 unverified; 8 pointer-only for licence)
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection 6 1 16 Nov 2023 ran 4 of 7 samples (3 unverified; 1 pointer-only for licence)
Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding 4 1 14 Nov 2023 not harvested
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models 2 1 8 Jun 2023 not harvested
VideoChat: Chat-Centric Video Understanding 1 1 10 May 2023 not harvested
HiTeA: Hierarchical Temporal-Aware Video-Language Pre-training 0 1 30 Dec 2022 not harvested
InternVideo: General Video Foundation Models via Generative and Discriminative Learning 2 1 6 Dec 2022 ran 3 of 3 samples (0 unverified)
Zero-Shot Video Question Answering via Frozen Bidirectional Language Models 3 1 16 Jun 2022 ran 14 of 34 samples (20 unverified; 1 pointer-only for licence)

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom (research-only)

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • TGIF-QA

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections