Datasets › TVQA

TVQA

Introduced by Jie Lei et al. in TVQA: Localized, Compositional Video Question Answering1 Jan 2018 archive 2025-07-28

The TVQA dataset is a large-scale video dataset for video question answering. It is based on 6 popular TV shows (Friends, The Big Bang Theory, How I Met Your Mother, House M.D., Grey's Anatomy, Castle). It includes 152,545 QA pairs from 21,793 TV show clips. The QA pairs are split into the ratio of 8:1:1 for training, validation, and test sets. The TVQA dataset provides the sequence of video frames extracted at 3 FPS, the corresponding subtitles with the video clips, and the query consisting of a question and four answer candidates. Among the four answer candidates, there is only one correct answer.

Source: Two-stream Spatiotemporal Feature for Video QA Task Image Source: https://arxiv.org/abs/1809.01696

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

11 shown of 11 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 146. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens 2 1 4 Apr 2024 ran 2 of 2 samples (0 unverified)
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM 1 1 27 Mar 2024 ran 3 of 3 samples (0 unverified)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark 3 4 28 Nov 2023 ran 7 of 10 samples (3 unverified)
Large Language Models are Temporal and Causal Reasoners for Video Question Answering 1 1 24 Oct 2023 ran 0 of 4 samples (4 unverified)
Self-Chained Image-Language Model for Video Localization and Question Answering 1 1 11 May 2023 ran 6 of 7 samples (1 unverified; 7 pointer-only for licence)
VindLU: A Recipe for Effective Video-and-Language Pretraining 1 1 9 Dec 2022 not harvested
InternVideo: General Video Foundation Models via Generative and Discriminative Learning 2 1 6 Dec 2022 ran 3 of 3 samples (0 unverified)
Zero-Shot Video Question Answering via Frozen Bidirectional Language Models 3 3 16 Jun 2022 ran 14 of 34 samples (20 unverified; 1 pointer-only for licence)
iPerceive: Applying Common-Sense Reasoning to Multi-Modal Dense Video Captioning and Video Question Answering 0 1 16 Nov 2020 not harvested
HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training 3 1 1 May 2020 ran 5 of 13 samples (8 unverified; 8 pointer-only for licence)
TVQA+: Spatio-Temporal Grounding for Video Question Answering 3 1 25 Apr 2019 ran 6 of 13 samples (7 unverified)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • TVQA

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections