Papers › Multiview Transformers for Video Recognition

Multiview Transformers for Video Recognition

12 Jan 2022CVPR 2022 1arXiv:2201.04288archive 2025-07-28

Shen Yan, Xuehan Xiong, Anurag Arnab, Zhichao Lu, Mi Zhang, Chen Sun, Cordelia Schmid

Video understanding requires reasoning at multiple spatiotemporal resolutions -- from short fine-grained motions to events taking place over longer durations. Although transformer architectures have recently advanced the state-of-the-art, they have not explicitly modelled different spatiotemporal resolutions. To this end, we present Multiview Transformers for Video Recognition (MTV). Our model consists of separate encoders to represent different views of the input video with lateral connections to fuse information across views. We present thorough ablation studies of our model and show that MTV consistently performs better than single-view counterparts in terms of accuracy and computational cost across a range of model sizes. Furthermore, we achieve state-of-the-art results on six standard datasets, and improve even further with large-scale pretraining. Code and checkpoints are available at: https://github.com/google-research/scenic/tree/main/scenic/projects/mtv.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

google-research/scenic officialmentioned in papermentioned on GitHubjax report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action ClassificationAction RecognitionVideo Understanding

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Classification Kinetics-400 MTV-H (WTS 60M) Acc@1 89.9 #14 of 207 Archive leaderboard report
Action Classification Kinetics-400 MTV-H (WTS 60M) Acc@5 98.3 #14 of 207 Archive leaderboard report
Action Classification Kinetics-400 MTV-H (WTS 60M) FLOPs (G) x views 735700x4x3 #14 of 207 Archive leaderboard report
Action Classification Kinetics-600 MTV-H (WTS 60M) Top-1 Accuracy 90.3 #9 of 65 Archive leaderboard report
Action Classification Kinetics-600 MTV-H (WTS 60M) Top-5 Accuracy 98.5 #9 of 65 Archive leaderboard report
Action Classification Kinetics-700 MTV-H (WTS 60M) Top-1 Accuracy 83.4 #6 of 36 Archive leaderboard report
Action Classification Kinetics-700 MTV-H (WTS 60M) Top-5 Accuracy 96.2 #6 of 36 Archive leaderboard report
Action Classification MiT MTV-H (WTS 60M) Top 1 Accuracy 47.2 #5 of 29 Archive leaderboard report
Action Classification MiT MTV-H (WTS 60M) Top 5 Accuracy 75.7 #5 of 29 Archive leaderboard report
Action Recognition EPIC-KITCHENS-100 MTV-B (WTS 60M) Action@1 50.5 #8 of 32 Archive leaderboard report
Action Recognition EPIC-KITCHENS-100 MTV-B (WTS 60M) Noun@1 63.9 #8 of 32 Archive leaderboard report
Action Recognition EPIC-KITCHENS-100 MTV-B (WTS 60M) Verb@1 69.9 #8 of 32 Archive leaderboard report
Action Recognition Something-Something V2 MTV-B Top-1 Accuracy 68.5 #50 of 123 Archive leaderboard report
Action Recognition Something-Something V2 MTV-B Top-5 Accuracy 90.4 #50 of 123 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections