Papers › Multimodal Open-Vocabulary Video Classification via Pre-Trained Vision and Language Models

Multimodal Open-Vocabulary Video Classification via Pre-Trained Vision and Language Models

15 Jul 2022arXiv:2207.07646archive 2025-07-28

Rui Qian, Yeqing Li, Zheng Xu, Ming-Hsuan Yang, Serge Belongie, Yin Cui

Utilizing vision and language models (VLMs) pre-trained on large-scale image-text pairs is becoming a promising paradigm for open-vocabulary visual recognition. In this work, we extend this paradigm by leveraging motion and audio that naturally exist in video. We present \textbf{MOV}, a simple yet effective method for \textbf{M}ultimodal \textbf{O}pen-\textbf{V}ocabulary video classification. In MOV, we directly use the vision encoder from pre-trained VLMs with minimal modifications to encode video, optical flow and audio spectrogram. We design a cross-modal fusion mechanism to aggregate complimentary multimodal information. Experiments on Kinetics-700 and VGGSound show that introducing flow or audio modality brings large performance gains over the pre-trained VLM and existing methods. Specifically, MOV greatly improves the accuracy on base classes, while generalizes better on novel classes. MOV achieves state-of-the-art results on UCF and HMDB zero-shot video classification benchmarks, significantly outperforming both traditional zero-shot methods and recent methods based on VLMs. Code and models will be released.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Optical Flow EstimationVideo ClassificationZero-Shot Action Recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Zero-Shot Action Recognition HMDB51 MOV (ViT-L/14) Top-1 Accuracy 64.7 #1 of 29 Archive leaderboard report
Zero-Shot Action Recognition HMDB51 MOV (ViT-B/16) Top-1 Accuracy 60.8 #4 of 29 Archive leaderboard report
Zero-Shot Action Recognition UCF101 MOV (ViT-L/14) Top-1 Accuracy 87.1 #3 of 35 Archive leaderboard report
Zero-Shot Action Recognition UCF101 MOV (ViT-B/16) Top-1 Accuracy 82.6 #9 of 35 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

BASE

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections