Papers › IntentQA: Context-aware Video Intent Reasoning
IntentQA: Context-aware Video Intent Reasoning
Jiapeng Li, Ping Wei, Wenjuan Han, Lifeng Fan
In this paper, we propose a novel task IntentQA, a special VideoQA task focusing on video intent reasoning, which has become increasingly important for AI with its advantages in equipping AI agents with the capability of reasoning beyond mere recognition in daily tasks. We also contribute a large-scale VideoQA dataset for this task. We propose a Context-aware Video Intent Reasoning model (CaVIR) consisting of i) Video Query Language (VQL) for better cross-modal representation of the situational context, ii) Contrastive Learning module for utilizing the contrastive context, and iii) Commonsense Reasoning module for incorporating the commonsense context. Comprehensive experiments on this challenging task demonstrate the effectiveness of each model component, the superiority of our full model over other baselines, and the generalizability of our model to a new VideoQA task. The dataset and codes are open-sourced at: https://github.com/JoseponLee/IntentQA.git
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Datasets
Introduced by this paper, per the archive.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Video Question Answering | IntentQA | Human | Accuarcy | 78.5 | #3 of 6 | Archive leaderboard | report |
| Video Question Answering | IntentQA | Human | CH | 80.2 | #3 of 6 | Archive leaderboard | report |
| Video Question Answering | IntentQA | Human | CW | 77.8 | #3 of 6 | Archive leaderboard | report |
| Video Question Answering | IntentQA | Human | TP&TN | 79.1 | #3 of 6 | Archive leaderboard | report |
| Video Question Answering | IntentQA | IntentQA | Accuarcy | 57.6 | #4 of 6 | Archive leaderboard | report |
| Video Question Answering | IntentQA | IntentQA | CH | 65.5 | #4 of 6 | Archive leaderboard | report |
| Video Question Answering | IntentQA | IntentQA | CW | 58.4 | #4 of 6 | Archive leaderboard | report |
| Video Question Answering | IntentQA | IntentQA | TP&TN | 50.5 | #4 of 6 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections