Papers › ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos

22 Nov 2024CVPR 2025 1arXiv:2411.14901archive 2025-07-28

Tanveer Hannan, Md Mohaiminul Islam, Jindong Gu, Thomas Seidl, Gedas Bertasius

Large language models (LLMs) excel at retrieving information from lengthy text, but their vision-language counterparts (VLMs) face difficulties with hour-long videos, especially for temporal grounding. Specifically, these VLMs are constrained by frame limitations, often losing essential temporal details needed for accurate event localization in extended video content. We propose ReVisionLLM, a recursive vision-language model designed to locate events in hour-long videos. Inspired by human search strategies, our model initially targets broad segments of interest, progressively revising its focus to pinpoint exact temporal boundaries. Our model can seamlessly handle videos of vastly different lengths, from minutes to hours. We also introduce a hierarchical training strategy that starts with short clips to capture distinct events and progressively extends to longer videos. To our knowledge, ReVisionLLM is the first VLM capable of temporal grounding in hour-long videos, outperforming previous state-of-the-art methods across multiple datasets by a significant margin (+2.6% R1@0.1 on MAD). The code is available at https://github.com/Tanveer81/ReVisionLLM.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

tanveer81/revisionllm officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Language ModelingLanguage ModellingLanguage-Based Temporal LocalizationNatural Language Moment Retrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Language-Based Temporal Localization VidChapters-7M ReVisionLLM R1@.9 15.2 #1 of 2 Archive leaderboard report
Natural Language Moment Retrieval MAD ReVisionLLM R@1,IoU=0.1 17.3 #1 of 8 Archive leaderboard report
Natural Language Moment Retrieval MAD ReVisionLLM R@1,IoU=0.3 12.7 #1 of 8 Archive leaderboard report
Natural Language Moment Retrieval MAD ReVisionLLM R@1,IoU=0.5 6.7 #1 of 8 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Focus

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections