Papers › WildQA: In-the-Wild Video Question Answering

WildQA: In-the-Wild Video Question Answering

14 Sep 2022arXiv:2209.06650archive 2025-07-28

Santiago Castro, Naihao Deng, Pingxuan Huang, Mihai Burzo, Rada Mihalcea

Existing video understanding datasets mostly focus on human interactions, with little attention being paid to the "in the wild" settings, where the videos are recorded outdoors. We propose WILDQA, a video understanding dataset of videos recorded in outside settings. In addition to video question answering (Video QA), we also introduce the new task of identifying visual support for a given question and answer (Video Evidence Selection). Through evaluations using a wide range of baseline models, we show that WILDQA poses new challenges to the vision and language research communities. The dataset is available at https://lit.eecs.umich.edu/wildqa/.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Evidence SelectionQuestion AnsweringVideo Question AnsweringVideo Understanding

Datasets

Introduced by this paper, per the archive.

WildQA

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Question Answering WildQA Multi (text + video, IO) ROUGE-1 34.0 ± 0.5 #1 of 5 Archive leaderboard report
Video Question Answering WildQA Multi (text + video, IO) ROUGE-2 18.8 ± 0.7 #1 of 5 Archive leaderboard report
Video Question Answering WildQA Multi (text + video, IO) ROUGE-L 32.8 ± 0.6 #1 of 5 Archive leaderboard report
Video Question Answering WildQA Multi (text + video, SE) ROUGE-1 33.8 ± 0.8 #2 of 5 Archive leaderboard report
Video Question Answering WildQA Multi (text + video, SE) ROUGE-2 18.5 ± 0.7 #2 of 5 Archive leaderboard report
Video Question Answering WildQA Multi (text + video, SE) ROUGE-L 32.5 ± 0.8 #2 of 5 Archive leaderboard report
Video Question Answering WildQA T5 (text) ROUGE-1 33.8 ± 0.2 #3 of 5 Archive leaderboard report
Video Question Answering WildQA T5 (text) ROUGE-2 17.7 ± 0.1 #3 of 5 Archive leaderboard report
Video Question Answering WildQA T5 (text) ROUGE-L 32.4 ± 0.3 #3 of 5 Archive leaderboard report
Video Question Answering WildQA T5 (text + video) ROUGE-1 33.1 ± 0.3 #4 of 5 Archive leaderboard report
Video Question Answering WildQA T5 (text + video) ROUGE-2 17.3 ± 0.4 #4 of 5 Archive leaderboard report
Video Question Answering WildQA T5 (text + video) ROUGE-L 31.9 ± 0.2 #4 of 5 Archive leaderboard report
Video Question Answering WildQA T5 (text, zero-shot) ROUGE-1 0.8 ± 0.0 #5 of 5 Archive leaderboard report
Video Question Answering WildQA T5 (text, zero-shot) ROUGE-2 0.0 ± 0.0 #5 of 5 Archive leaderboard report
Video Question Answering WildQA T5 (text, zero-shot) ROUGE-L 0.8 ± 0.0 #5 of 5 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections