Papers › Video Transformer Network
Video Transformer Network
Daniel Neimark, Omri Bar, Maya Zohar, Dotan Asselmann
This paper presents VTN, a transformer-based framework for video recognition. Inspired by recent developments in vision transformers, we ditch the standard approach in video action recognition that relies on 3D ConvNets and introduce a method that classifies actions by attending to the entire video sequence information. Our approach is generic and builds on top of any given 2D spatial network. In terms of wall runtime, it trains 16.1× faster and runs 5.1× faster during inference while maintaining competitive accuracy compared to other state-of-the-art methods. It enables whole video analysis, via a single end-to-end pass, while requiring 1.5× fewer GFLOPs. We report competitive results on Kinetics-400 and present an ablation study of VTN properties and the trade-off between accuracy and inference speed. We hope our approach will serve as a new baseline and start a fresh line of research in the video recognition domain. Code and models are available at: https://github.com/bomri/SlowFast/blob/master/projects/vtn/README.md
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Action Classification | Kinetics-400 | ViT-B-VTN+ ImageNet-21K (84.0 [10]) | Acc@1 | 79.8 | #107 of 207 | Archive leaderboard | report |
| Action Classification | Kinetics-400 | ViT-B-VTN (3 layers, ImageNet pretrain) | Acc@1 | 78.6 | #125 of 207 | Archive leaderboard | report |
| Action Classification | Kinetics-400 | ViT-B-VTN (3 layers, ImageNet pretrain) | Acc@5 | 93.7 | #125 of 207 | Archive leaderboard | report |
| Action Classification | Kinetics-400 | ViT-B-VTN+ ImageNet-21K (84.0 [10]) | Acc@5 | 94.2 | #204 of 207 | Archive leaderboard | report |
| Action Classification | Kinetics-400 | ViT-B-VTN (1 layer, ImageNet pretrain) | Acc@5 | 93.4 | #206 of 207 | Archive leaderboard | report |
| Action Classification | MiT | VTN | Top 1 Accuracy | 37.4 | #12 of 29 | Archive leaderboard | report |
| Action Classification | MiT | VTN | Top 5 Accuracy | 65.4 | #12 of 29 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections