Papers › iPerceive: Applying Common-Sense Reasoning to Multi-Modal Dense Video Captioning and...

iPerceive: Applying Common-Sense Reasoning to Multi-Modal Dense Video Captioning and Video Question Answering

16 Nov 2020arXiv:2011.07735archive 2025-07-28

Aman Chadha, Gurneet Arora, Navpreet Kaloty

Most prior art in visual understanding relies solely on analyzing the "what" (e.g., event recognition) and "where" (e.g., event localization), which in some cases, fails to describe correct contextual relationships between events or leads to incorrect underlying visual attention. Part of what defines us as human and fundamentally different from machines is our instinct to seek causality behind any association, say an event Y that happened as a direct result of event X. To this end, we propose iPerceive, a framework capable of understanding the "why" between events in a video by building a common-sense knowledge base using contextual cues to infer causal relationships between objects in the video. We demonstrate the effectiveness of our technique using the dense video captioning (DVC) and video question answering (VideoQA) tasks. Furthermore, while most prior work in DVC and VideoQA relies solely on visual information, other modalities such as audio and speech are vital for a human observer's perception of an environment. We formulate DVC and VideoQA tasks as machine translation problems that utilize multiple modalities. By evaluating the performance of iPerceive DVC and iPerceive VideoQA on the ActivityNet Captions and TVQA datasets respectively, we show that our approach furthers the state-of-the-art. Code and samples are available at: iperceive.amanchadha.com.

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Common Sense ReasoningDense Video CaptioningMachine TranslationQuestion AnsweringVideo CaptioningVideo Question Answering

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Dense Video Captioning ActivityNet Captions iPerceive (Chadha et al., 2020) BLEU-3 2.93 #10 of 12 Archive leaderboard report
Dense Video Captioning ActivityNet Captions iPerceive (Chadha et al., 2020) BLEU-4 1.29 #10 of 12 Archive leaderboard report
Dense Video Captioning ActivityNet Captions iPerceive (Chadha et al., 2020) METEOR 7.87 #10 of 12 Archive leaderboard report
Video Question Answering TVQA iPerceive (Chadha et al., 2020) Accuracy 76.96 #4 of 6 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections