Papers › Bidirectional Cross-Modal Knowledge Exploration for Video Recognition with Pre-trained...

Bidirectional Cross-Modal Knowledge Exploration for Video Recognition with Pre-trained Vision-Language Models

31 Dec 2022CVPR 2023 1arXiv:2301.00182archive 2025-07-28

Wenhao Wu, Xiaohan Wang, Haipeng Luo, Jingdong Wang, Yi Yang, Wanli Ouyang

Vision-language models (VLMs) pre-trained on large-scale image-text pairs have demonstrated impressive transferability on various visual tasks. Transferring knowledge from such powerful VLMs is a promising direction for building effective video recognition models. However, current exploration in this field is still limited. We believe that the greatest value of pre-trained VLMs lies in building a bridge between visual and textual domains. In this paper, we propose a novel framework called BIKE, which utilizes the cross-modal bridge to explore bidirectional knowledge: i) We introduce the Video Attribute Association mechanism, which leverages the Video-to-Text knowledge to generate textual auxiliary attributes for complementing video recognition. ii) We also present a Temporal Concept Spotting mechanism that uses the Text-to-Video expertise to capture temporal saliency in a parameter-free manner, leading to enhanced video representation. Extensive studies on six popular video datasets, including Kinetics-400 & 600, UCF-101, HMDB-51, ActivityNet and Charades, show that our method achieves state-of-the-art performance in various recognition scenarios, such as general, zero-shot, and few-shot video recognition. Our best model achieves a state-of-the-art accuracy of 88.6% on the challenging Kinetics-400 using the released CLIP model. The code is available at https://github.com/whwu95/BIKE .

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

whwu95/BIKE officialmentioned in papermentioned on GitHubpytorchMIT report
whwu95/ATM mentioned on GitHubpytorch report
whwu95/Cap4Video mentioned on GitHubpytorchMIT report
whwu95/GPT4Vis mentioned on GitHubMIT report
whwu95/text4vis mentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action ClassificationAction RecognitionAttributeVideo RecognitionZero-Shot Action Recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Classification Charades BIKE MAP 50.7 #10 of 49 Archive leaderboard report
Action Classification Kinetics-400 BIKE (CLIP ViT-L/14) Acc@1 88.7 #21 of 207 Archive leaderboard report
Action Classification Kinetics-400 BIKE (CLIP ViT-L/14) Acc@5 98.4 #21 of 207 Archive leaderboard report
Action Recognition ActivityNet BIKE mAP 96.1 #2 of 16 Archive leaderboard report
Action Recognition HMDB-51 BIKE Average accuracy of 3 splits 83.1 #13 of 77 Archive leaderboard report
Action Recognition UCF101 BIKE 3-fold Accuracy 98.8 #5 of 91 Archive leaderboard report
Zero-Shot Action Recognition ActivityNet BIKE Top-1 Accuracy 86.2 #1 of 5 Archive leaderboard report
Zero-Shot Action Recognition HMDB51 BIKE Top-1 Accuracy 61.4 #3 of 29 Archive leaderboard report
Zero-Shot Action Recognition Kinetics BIKE Top-1 Accuracy 68.5 #8 of 20 Archive leaderboard report
Zero-Shot Action Recognition Kinetics BIKE Top-5 Accuracy 91.1 #8 of 20 Archive leaderboard report
Zero-Shot Action Recognition UCF101 BIKE Top-1 Accuracy 86.6 #5 of 35 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

CLIP

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections