Papers › Just Ask: Learning to Answer Questions from Millions of Narrated Videos

Just Ask: Learning to Answer Questions from Millions of Narrated Videos

1 Dec 2020ICCV 2021 10arXiv:2012.00451archive 2025-07-28

Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, Cordelia Schmid

Recent methods for visual question answering rely on large-scale annotated datasets. Manual annotation of questions and answers for videos, however, is tedious, expensive and prevents scalability. In this work, we propose to avoid manual annotation and generate a large-scale training dataset for video question answering making use of automatic cross-modal supervision. We leverage a question generation transformer trained on text data and use it to generate question-answer pairs from transcribed video narrations. Given narrated videos, we then automatically generate the HowToVQA69M dataset with 69M video-question-answer triplets. To handle the open vocabulary of diverse answers in this dataset, we propose a training procedure based on a contrastive loss between a video-question multi-modal transformer and an answer transformer. We introduce the zero-shot VideoQA task and show excellent results, in particular for rare answers. Furthermore, we demonstrate our method to significantly outperform the state of the art on MSRVTT-QA, MSVD-QA, ActivityNet-QA and How2QA. Finally, for a detailed evaluation we introduce iVQA, a new VideoQA dataset with reduced language biases and high-quality redundant manual annotations. Our code, datasets and trained models are available at https://antoyang.github.io/just-ask.html.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

antoyang/just-ask officialmentioned on GitHubpytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Question AnsweringQuestion GenerationQuestion-GenerationVideo Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Zero-Shot Learning

Datasets

Introduced by this paper, per the archive.

HowToVQA69MiVQA

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Question Answering ActivityNet-QA Just Ask (fine-tune) Accuracy 38.9 #26 of 36 Archive leaderboard report
Video Question Answering ActivityNet-QA Just Ask (0-shot) Accuracy 12.2 #36 of 36 Archive leaderboard report
Video Question Answering How2QA Just Ask Accuracy 84.4 #3 of 8 Archive leaderboard report
Video Question Answering How2QA Just Ask (0-shot) Accuracy 51.1 #8 of 8 Archive leaderboard report
Video Question Answering VideoQA Just Ask (fine-tune) Accuracy 15.6 #1 of 1 Archive leaderboard report
Video Question Answering iVQA Just Ask (fine-tune) Accuracy 35.4 #5 of 7 Archive leaderboard report
Video Question Answering iVQA Just Ask (0-shot) Accuracy 12.2 #7 of 7 Archive leaderboard report
Visual Question Answering MSRVTT-QA Just Ask Accuracy 0.415 #3 of 4 Archive leaderboard report
Visual Question Answering MSVD-QA Just Ask Accuracy 0.463 #2 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections