Papers › Video Transformer Network

Video Transformer Network

1 Feb 2021arXiv:2102.00719archive 2025-07-28

Daniel Neimark, Omri Bar, Maya Zohar, Dotan Asselmann

This paper presents VTN, a transformer-based framework for video recognition. Inspired by recent developments in vision transformers, we ditch the standard approach in video action recognition that relies on 3D ConvNets and introduce a method that classifies actions by attending to the entire video sequence information. Our approach is generic and builds on top of any given 2D spatial network. In terms of wall runtime, it trains 16.1× faster and runs 5.1× faster during inference while maintaining competitive accuracy compared to other state-of-the-art methods. It enables whole video analysis, via a single end-to-end pass, while requiring 1.5× fewer GFLOPs. We report competitive results on Kinetics-400 and present an ablation study of VTN properties and the trade-off between accuracy and inference speed. We hope our approach will serve as a new baseline and start a fresh line of research in the video recognition domain. Code and models are available at: https://github.com/bomri/SlowFast/blob/master/projects/vtn/README.md

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

bomri/SlowFast officialpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action ClassificationAction RecognitionAction Recognition In VideosVideo Recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Classification Kinetics-400 ViT-B-VTN+ ImageNet-21K (84.0 [10]) Acc@1 79.8 #107 of 207 Archive leaderboard report
Action Classification Kinetics-400 ViT-B-VTN (3 layers, ImageNet pretrain) Acc@1 78.6 #125 of 207 Archive leaderboard report
Action Classification Kinetics-400 ViT-B-VTN (3 layers, ImageNet pretrain) Acc@5 93.7 #125 of 207 Archive leaderboard report
Action Classification Kinetics-400 ViT-B-VTN+ ImageNet-21K (84.0 [10]) Acc@5 94.2 #204 of 207 Archive leaderboard report
Action Classification Kinetics-400 ViT-B-VTN (1 layer, ImageNet pretrain) Acc@5 93.4 #206 of 207 Archive leaderboard report
Action Classification MiT VTN Top 1 Accuracy 37.4 #12 of 29 Archive leaderboard report
Action Classification MiT VTN Top 5 Accuracy 65.4 #12 of 29 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections