{"url":"/dataset/intentqa","name":"IntentQA","full_name":null,"description_markdown":"We contribute an IntentQA dataset with diverse intents in daily social activities.\r\n\r\nWe utilize NExT-QA as the source dataset to construct our dataset. NExT-QA dataset is a comprehensive VideoQA dataset with rich natural daily social activities and detailed QA annotations. Originally, the NExT-QA dataset categorizes itself into three types, i.e., Causal, Temporal, Descriptive. We select the inference QA types, i.e., Causal and Temporal, rather than the factoid Descriptive, to build our IntentQA dataset. Particularly, we select both the Causal Why and Causal How subtypes under Causal, and the Temporal Previous and Temporal Next subtypes under Temporal. The Causal Why (CW) QA usually takes the form of ‘Why [action]? For [intent]’, with the key action appearing in the question and the intent in the answer. On the contrary, the Causal How (CH) QA usually takes the form of ‘How [intent]? By [action]’, with the key action appearing in the answer and the intent in the question. The Temporal Previous (TP) QA usually takes the form of ‘What [action A] before [action B]? ’, while the Temporal Next (TN) QA takes the form of ‘What [action B] after [action A]? ’. In the TP&TN QA, the intent is not explicitly expressed in the question nor answer, but is the implicit causal factor linking the two sequential actions.","description_withheld":null,"homepage":"https://github.com/JoseponLee/IntentQA","introduced_date":"2023-08-16","introduced_date_note":null,"introduced_by":{"paper":"/paper/intentqa-context-aware-video-intent-reasoning","title":"IntentQA: Context-aware Video Intent Reasoning","first_author":"Jiapeng Li","url":null},"license":null,"modalities":[{"name":"Videos","url":"/datasets/modality/videos"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Video Question Answering","url":"/task/video-question-answering","datasets_with_task":"/datasets/task/video-question-answering"},{"name":"Zero-Shot Video Question Answer","url":"/task/zeroshot-video-question-answer","datasets_with_task":"/datasets/task/zeroshot-video-question-answer"}],"languages":[],"variants":["IntentQA"],"data_loaders":[],"num_papers_in_archive":27,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/zero-shot-video-question-answer-on-intentqa","task":"Zero-Shot Video Question Answer","dataset_variant":"IntentQA","rows":13,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"ENTER","paper":"/paper/enter-event-based-interpretable-reasoning-for","metrics":{"Accuracy":"71.5"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/video-question-answering-on-intentqa","task":"Video Question Answering","dataset_variant":"IntentQA","rows":6,"metrics":["Accuarcy","CW","CH","TP&TN"],"first_row_in_archive_order":{"model":"VideoChat2_HD_mistral","paper":"/paper/mvbench-a-comprehensive-multi-modal-video","metrics":{"Accuarcy":"83.4","CH":"90.0","CW":"84.0","TP&TN":"77.3"},"code_links":[{"title":"opengvlab/ask-anything","url":"https://github.com/opengvlab/ask-anything"},{"title":"magic-research/PLLaVA","url":"https://github.com/magic-research/PLLaVA"},{"title":"bytedance/tarsier","url":"https://github.com/bytedance/tarsier"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/enter-event-based-interpretable-reasoning-for","title":"ENTER: Event Based Interpretable Reasoning for VideoQA","date":"2025-01-24","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/vidctx-context-aware-video-question-answering","title":"VidCtx: Context-aware Video Question Answering with Image Models","date":"2024-12-23","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/ts-llava-constructing-visual-tokens-through","title":"TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models","date":"2024-11-17","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":11,"samples_ran":5,"samples_unverified":6,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/slowfast-llava-a-strong-training-free","title":"SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models","date":"2024-07-22","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":4,"samples_unverified":2,"pointer_only_for_licence":6,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/too-many-frames-not-all-useful-efficient","title":"Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA","date":"2024-06-13","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":3,"samples_unverified":0,"pointer_only_for_licence":3,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/videotree-adaptive-tree-based-video","title":"VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos","date":"2024-05-29","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":11,"samples_ran":11,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/an-image-grid-can-be-worth-a-video-zero-shot","title":"An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM","date":"2024-03-27","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":3,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/language-repository-for-long-video","title":"Language Repository for Long Video Understanding","date":"2024-03-21","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":9,"samples_ran":9,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/a-simple-llm-framework-for-long-range-video","title":"A Simple LLM Framework for Long-Range Video Question-Answering","date":"2023-12-28","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":6,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/mvbench-a-comprehensive-multi-modal-video","title":"MVBench: A Comprehensive Multi-modal Video Understanding Benchmark","date":"2023-11-28","rows_on_this_dataset":2,"code_links":3,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":10,"samples_ran":7,"samples_unverified":3,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/mistral-7b","title":"Mistral 7B","date":"2023-10-10","rows_on_this_dataset":1,"code_links":6,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":11,"samples_ran":9,"samples_unverified":2,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/self-chained-image-language-model-for-video-1","title":"Self-Chained Image-Language Model for Video Localization and Question Answering","date":"2023-05-11","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":7,"samples_ran":6,"samples_unverified":1,"pointer_only_for_licence":7,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/intentqa-context-aware-video-intent-reasoning","title":"IntentQA: Context-aware Video Intent Reasoning","date":"2023-01-01","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/video-graph-transformer-for-video-question","title":"Video Graph Transformer for Video Question Answering","date":"2022-07-12","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":14,"samples_ran":9,"samples_unverified":5,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/video-as-conditional-graph-hierarchy-for","title":"Video as Conditional Graph Hierarchy for Multi-Granular Question Answering","date":"2021-12-12","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":1,"samples_unverified":5,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":12,"samples_harvested":97,"samples_ran":73,"samples_unverified":24,"pointer_only_for_licence":17,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}