Papers › Co-training Transformer with Videos and Images Improves Action Recognition

Co-training Transformer with Videos and Images Improves Action Recognition

14 Dec 2021arXiv:2112.07175archive 2025-07-28

BoWen Zhang, Jiahui Yu, Christopher Fifty, Wei Han, Andrew M. Dai, Ruoming Pang, Fei Sha

In learning action recognition, models are typically pre-trained on object recognition with images, such as ImageNet, and later fine-tuned on target action recognition with videos. This approach has achieved good empirical performance especially with recent transformer-based video architectures. While recently many works aim to design more advanced transformer architectures for action recognition, less effort has been made on how to train video transformers. In this work, we explore several training paradigms and present two findings. First, video transformers benefit from joint training on diverse video datasets and label spaces (e.g., Kinetics is appearance-focused while SomethingSomething is motion-focused). Second, by further co-training with images (as single-frame videos), the video transformers learn even better video representations. We term this approach as Co-training Videos and Images for Action Recognition (CoVeR). In particular, when pretrained on ImageNet-21K based on the TimeSFormer architecture, CoVeR improves Kinetics-400 Top-1 Accuracy by 2.4%, Kinetics-600 by 2.3%, and SomethingSomething-v2 by 2.3%. When pretrained on larger-scale image datasets following previous state-of-the-art, CoVeR achieves best results on Kinetics-400 (87.2%), Kinetics-600 (87.9%), Kinetics-700 (79.8%), SomethingSomething-v2 (70.9%), and Moments-in-Time (46.1%), with a simple spatio-temporal video transformer.

PaperPDF

In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action ClassificationAction RecognitionAction Recognition In VideosObject RecognitionVideo Classification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Classification Kinetics-400 CoVeR (JFT-3B) Acc@1 87.2 #39 of 207 Archive leaderboard report
Action Classification Kinetics-400 CoVeR (JFT-3B) Acc@5 97.5 #39 of 207 Archive leaderboard report
Action Classification Kinetics-400 CoVeR (JFT-300M) Acc@1 86.3 #49 of 207 Archive leaderboard report
Action Classification Kinetics-400 CoVeR (JFT-300M) Acc@5 97.2 #49 of 207 Archive leaderboard report
Action Classification Kinetics-600 CoVeR (JFT-3B) Top-1 Accuracy 87.9 #24 of 65 Archive leaderboard report
Action Classification Kinetics-600 CoVeR (JFT-3B) Top-5 Accuracy 97.8 #24 of 65 Archive leaderboard report
Action Classification Kinetics-600 CoVeR (JFT-300M) Top-1 Accuracy 86.8 #26 of 65 Archive leaderboard report
Action Classification Kinetics-600 CoVeR (JFT-300M) Top-5 Accuracy 97.3 #26 of 65 Archive leaderboard report
Action Classification Kinetics-700 CoVeR (JFT-3B) Top-1 Accuracy 79.8 #15 of 36 Archive leaderboard report
Action Classification Kinetics-700 CoVeR (JFT-3B) Top-5 Accuracy 94.9 #15 of 36 Archive leaderboard report
Action Classification Kinetics-700 CoVeR (JFT-300M) Top-1 Accuracy 78.5 #18 of 36 Archive leaderboard report
Action Classification Kinetics-700 CoVeR (JFT-300M) Top-5 Accuracy 94.2 #18 of 36 Archive leaderboard report
Action Classification MiT CoVeR(JFT-3B) Top 1 Accuracy 46.1 #6 of 29 Archive leaderboard report
Action Classification MiT CoVeR(JFT-3B) Top 5 Accuracy 75.4 #6 of 29 Archive leaderboard report
Action Classification MiT CoVeR(JFT-300M) Top 1 Accuracy 45.0 #7 of 29 Archive leaderboard report
Action Classification MiT CoVeR(JFT-300M) Top 5 Accuracy 73.9 #7 of 29 Archive leaderboard report
Action Recognition Something-Something V2 CoVeR(JFT-3B) Top-1 Accuracy 70.9 #35 of 123 Archive leaderboard report
Action Recognition Something-Something V2 CoVeR(JFT-3B) Top-5 Accuracy 92.5 #35 of 123 Archive leaderboard report
Action Recognition Something-Something V2 CoVeR(JFT-300M) Top-1 Accuracy 69.8 #40 of 123 Archive leaderboard report
Action Recognition Something-Something V2 CoVeR(JFT-300M) Top-5 Accuracy 91.9 #40 of 123 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections