Papers › RTQ: Rethinking Video-language Understanding Based on Image-text Model

RTQ: Rethinking Video-language Understanding Based on Image-text Model

1 Dec 2023arXiv:2312.00347archive 2025-07-28

Xiao Wang, Yaoyu Li, Tian Gan, Zheng Zhang, Jingjing Lv, Liqiang Nie

Recent advancements in video-language understanding have been established on the foundation of image-text models, resulting in promising outcomes due to the shared knowledge between images and videos. However, video-language understanding presents unique challenges due to the inclusion of highly complex semantic details, which result in information redundancy, temporal dependency, and scene complexity. Current techniques have only partially tackled these issues, and our quantitative analysis indicates that some of these methods are complementary. In light of this, we propose a novel framework called RTQ (Refine, Temporal model, and Query), which addresses these challenges simultaneously. The approach involves refining redundant information within frames, modeling temporal relations among frames, and querying task-specific information from the videos. Remarkably, our model demonstrates outstanding performance even in the absence of video-language pre-training, and the results are comparable with or superior to those achieved by state-of-the-art pre-training methods. Code is available at https://github.com/SCZwangxiao/RTQ-MM2023.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

SCZwangxiao/RTQ-MM2023 officialmentioned in papermentioned on GitHubpytorch report
sczwangxiao/tsgvs-mm2023 mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Video CaptioningVideo Question AnsweringVideo Retrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Captioning MSR-VTT RTQ BLEU-4 49.6 #9 of 24 Archive leaderboard report
Video Captioning MSR-VTT RTQ CIDEr 69.3 #9 of 24 Archive leaderboard report
Video Captioning MSR-VTT RTQ ROUGE-L 66.1 #9 of 24 Archive leaderboard report
Video Captioning MSVD RTQ BLEU-4 66.9 #10 of 16 Archive leaderboard report
Video Captioning MSVD RTQ CIDEr 123.4 #10 of 16 Archive leaderboard report
Video Captioning MSVD RTQ ROUGE-L 82.2 #10 of 16 Archive leaderboard report
Video Question Answering NExT-QA RTQ Accuracy 63.2 #32 of 47 Archive leaderboard report
Video Retrieval ActivityNet RTQ text-to-video R@1 53.5 #13 of 31 Archive leaderboard report
Video Retrieval ActivityNet RTQ text-to-video R@10 91.9 #13 of 31 Archive leaderboard report
Video Retrieval ActivityNet RTQ text-to-video R@5 81.4 #13 of 31 Archive leaderboard report
Video Retrieval DiDeMo RTQ text-to-video R@1 57.6 #11 of 40 Archive leaderboard report
Video Retrieval DiDeMo RTQ text-to-video R@10 89.9 #11 of 40 Archive leaderboard report
Video Retrieval DiDeMo RTQ text-to-video R@5 84.1 #11 of 40 Archive leaderboard report
Video Retrieval MSR-VTT-1kA RTQ text-to-video R@1 53.4 #10 of 63 Archive leaderboard report
Video Retrieval MSR-VTT-1kA RTQ text-to-video R@10 84.4 #10 of 63 Archive leaderboard report
Video Retrieval MSR-VTT-1kA RTQ text-to-video R@5 76.1 #10 of 63 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections