Papers › VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video...

VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval

2 Dec 2024arXiv:2412.01558archive 2025-07-28

Dhiman Paul, Md Rizwan Parvez, Nabeel Mohammed, Shafin Rahman

Video Highlight Detection and Moment Retrieval (HD/MR) are essential in video analysis. Recent joint prediction transformer models often overlook their cross-task dynamics and video-text alignment and refinement. Moreover, most models typically use limited, uni-directional attention mechanisms, resulting in weakly integrated representations and suboptimal performance in capturing the interdependence between video and text modalities. Although large-language and vision-language models (LLM/LVLMs) have gained prominence across various domains, their application in this field remains relatively underexplored. Here we propose VideoLights, a novel HD/MR framework addressing these limitations through (i) Convolutional Projection and Feature Refinement modules with an alignment loss for better video-text feature alignment, (ii) Bi-Directional Cross-Modal Fusion network for strongly coupled query-aware clip representations, and (iii) Uni-directional joint-task feedback mechanism enhancing both tasks through correlation. In addition, (iv) we introduce hard positive/negative losses for adaptive error penalization and improved learning, and (v) leverage LVLMs like BLIP-2 for enhanced multimodal feature integration and intelligent pretraining using synthetic data generated from LVLMs. Comprehensive experiments on QVHighlights, TVSum, and Charades-STA benchmarks demonstrate state-of-the-art performance. Codes and models are available at https://github.com/dpaul06/VideoLights .

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

dpaul06/VideoLights officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Highlight DetectionMoment Retrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Highlight Detection QVHighlights VideoLights-B-pt Hit@1 70.56 #4 of 21 Archive leaderboard report
Highlight Detection QVHighlights VideoLights-B-pt mAP 42.84 #4 of 21 Archive leaderboard report
Moment Retrieval Charades-STA VideoLights-B-pt R@1 IoU=0.3 73.33 #10 of 25 Archive leaderboard report
Moment Retrieval Charades-STA VideoLights-B-pt R@1 IoU=0.5 61.96 #10 of 25 Archive leaderboard report
Moment Retrieval Charades-STA VideoLights-B-pt R@1 IoU=0.7 41.05 #10 of 25 Archive leaderboard report
Moment Retrieval Charades-STA VideoLights-B-pt mIoU 52.94 #10 of 25 Archive leaderboard report
Moment Retrieval QVHighlights VideoLights-B-pt R@1 IoU=0.5 70.36 #7 of 32 Archive leaderboard report
Moment Retrieval QVHighlights VideoLights-B-pt R@1 IoU=0.7 55.25 #7 of 32 Archive leaderboard report
Moment Retrieval QVHighlights VideoLights-B-pt mAP 47.94 #7 of 32 Archive leaderboard report
Moment Retrieval QVHighlights VideoLights-B-pt mAP@0.5 69.53 #7 of 32 Archive leaderboard report
Moment Retrieval QVHighlights VideoLights-B-pt mAP@0.75 49.17 #7 of 32 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AttentionCLIPSoftmax

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections