Papers › Revisiting Temporal Modeling for CLIP-based Image-to-Video Knowledge Transferring

Revisiting Temporal Modeling for CLIP-based Image-to-Video Knowledge Transferring

26 Jan 2023CVPR 2023 1arXiv:2301.11116archive 2025-07-28

Ruyang Liu, Jingjia Huang, Ge Li, Jiashi Feng, Xinglong Wu, Thomas H. Li

Image-text pretrained models, e.g., CLIP, have shown impressive general multi-modal knowledge learned from large-scale image-text data pairs, thus attracting increasing attention for their potential to improve visual representation learning in the video domain. In this paper, based on the CLIP model, we revisit temporal modeling in the context of image-to-video knowledge transferring, which is the key point for extending image-text pretrained models to the video domain. We find that current temporal modeling mechanisms are tailored to either high-level semantic-dominant tasks (e.g., retrieval) or low-level visual pattern-dominant tasks (e.g., recognition), and fail to work on the two cases simultaneously. The key difficulty lies in modeling temporal dependency while taking advantage of both high-level and low-level knowledge in CLIP model. To tackle this problem, we present Spatial-Temporal Auxiliary Network (STAN) -- a simple and effective temporal modeling mechanism extending CLIP model to diverse video tasks. Specifically, to realize both low-level and high-level knowledge transferring, STAN adopts a branch structure with decomposed spatial-temporal modules that enable multi-level CLIP features to be spatial-temporally contextualized. We evaluate our method on two representative video tasks: Video-Text Retrieval and Video Recognition. Extensive experiments demonstrate the superiority of our model over the state-of-the-art methods on various datasets, including MSR-VTT, DiDeMo, LSMDC, MSVD, Kinetics-400, and Something-Something-V2. Codes will be available at https://github.com/farewellthree/STAN

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

farewellthree/stan officialmentioned in papermentioned on GitHubpytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Representation LearningRetrievalText RetrievalVideo RecognitionVideo RetrievalVideo-Text Retrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Retrieval DiDeMo STAN text-to-video Median Rank 1 #17 of 40 Archive leaderboard report
Video Retrieval DiDeMo STAN text-to-video R@1 54.6 #17 of 40 Archive leaderboard report
Video Retrieval DiDeMo STAN text-to-video R@10 85.1 #17 of 40 Archive leaderboard report
Video Retrieval DiDeMo STAN text-to-video R@5 78.4 #17 of 40 Archive leaderboard report
Video Retrieval LSMDC STAN text-to-video Median Rank 6 #11 of 38 Archive leaderboard report
Video Retrieval LSMDC STAN text-to-video R@1 29.2 #11 of 38 Archive leaderboard report
Video Retrieval LSMDC STAN text-to-video R@10 58.8 #11 of 38 Archive leaderboard report
Video Retrieval LSMDC STAN text-to-video R@5 49.5 #11 of 38 Archive leaderboard report
Video Retrieval MSR-VTT-1kA STAN text-to-video Median Rank 1 #7 of 63 Archive leaderboard report
Video Retrieval MSR-VTT-1kA STAN text-to-video R@1 54.1 #7 of 63 Archive leaderboard report
Video Retrieval MSR-VTT-1kA STAN text-to-video R@10 87.8 #7 of 63 Archive leaderboard report
Video Retrieval MSR-VTT-1kA STAN text-to-video R@5 79.5 #7 of 63 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

CLIPfail

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections