{"url":"/dataset/implicitqa","name":"ImplicitQA","full_name":null,"description_markdown":"The ImplicitQA dataset was introduced in the paper [ImplicitQA: Going beyond frames towards Implicit Video Reasoning](https://arxiv.org/abs/2506.21742).\r\n\r\n**Project page:** https://swetha5.github.io/ImplicitQA/\r\n\r\nImplicitQA is a novel benchmark specifically designed to test models on implicit reasoning in Video Question Answering (VideoQA). Unlike existing VideoQA benchmarks that primarily focus on questions answerable through explicit visual content (actions, objects, events directly observable within individual frames or short clips), ImplicitQA addresses the need for models to infer motives, causality, and relationships <u>across discontinuous frames</u>. This mirrors human-like understanding of creative and cinematic videos, which often employ storytelling techniques that deliberately omit certain depictions.\r\n\r\nThe dataset comprises 1,000 meticulously annotated QA pairs derived from over 320 high-quality creative video clips. These QA pairs are systematically categorized into key reasoning dimensions, including:\r\n\r\n*   Lateral spatial reasoning\r\n*   Vertical spatial reasoning\r\n*   Relative Depth and proximity\r\n*   Viewpoint and visibility\r\n*   Motion and trajectory Dynamics\r\n*   Causal and motivational reasoning\r\n*   Social interactions and Relationships\r\n*   Physical and Environmental context\r\n*   Inferred counting\r\n\r\nThe annotations are deliberately challenging, crafted to ensure high quality and to highlight the difficulty of implicit reasoning for current VideoQA models.","description_withheld":null,"homepage":"https://huggingface.co/datasets/ucf-crcv/ImplicitQA","introduced_date":"2025-05-14","introduced_date_note":null,"introduced_by":{"paper":"/paper/implicitqa-going-beyond-frames-towards","title":"ImplicitQA: Going beyond frames towards Implicit Video Reasoning","first_author":"Sirnam Swetha","url":null},"license":null,"modalities":[{"name":"Videos","url":"/datasets/modality/videos"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"task","url":null,"datasets_with_task":"/datasets/task/task"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["ImplicitQA"],"data_loaders":[],"num_papers_in_archive":5,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/on-implicitqa","task":"","dataset_variant":"ImplicitQA","rows":7,"metrics":["Average Accuracy","Macro Average Accuracy"],"first_row_in_archive_order":{"model":"GPT O3","paper":"/paper/implicitqa-going-beyond-frames-towards","metrics":{"Average Accuracy":"64.1","Macro Average Accuracy":"68.6"},"code_links":[{"title":"UCF-CRCV/ImplicitQA","url":"https://github.com/UCF-CRCV/ImplicitQA"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/implicitqa-going-beyond-frames-towards","title":"ImplicitQA: Going beyond frames towards Implicit Video Reasoning","date":"2025-06-26","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/qwen2-5-vl-technical-report","title":"Qwen2.5-VL Technical Report","date":"2025-02-19","rows_on_this_dataset":1,"code_links":4,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":2,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/video-instruction-tuning-with-synthetic-data","title":"Video Instruction Tuning With Synthetic Data","date":"2024-10-03","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/qwen2-vl-enhancing-vision-language-model-s","title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution","date":"2024-09-18","rows_on_this_dataset":1,"code_links":8,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":12,"samples_ran":8,"samples_unverified":4,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/llava-onevision-easy-visual-task-transfer","title":"LLaVA-OneVision: Easy Visual Task Transfer","date":"2024-08-06","rows_on_this_dataset":1,"code_links":2,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":2,"samples_harvested":15,"samples_ran":10,"samples_unverified":5,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}