Browse State-of-the-Art › Video Classification
Video Classification
206 papers with code · 12 benchmarks · 22 datasets archive 2025-07-28
Video Classification is the task of producing a label that is relevant to the video given its frames. A good video level classifier is one that not only provides accurate frame labels, but also best describes the entire video given the features and the annotations of the various frames in the video. For example, a video might contain a tree in some frame, but the label that is central to the video might be something else (e.g., “hiking”). The granularity of the labels that are needed to describe the frames and the video depends on the task. Typical tasks include assigning one or more global labels to the video, and assigning one or more labels for each frame inside the video.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
12 leaderboard tables shown for this task, 12 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 12 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
22 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
2 subtasks in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 206 papers with code (455 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
21 Nov 2017 32 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)Both convolutional and recurrent operations are building blocks that process one local neighborhood at a time.
-
22 Mar 2018 22 repositories listed Syntology ran 5 of 15 samples · 10 unverified · 4 pointer-only (licence)FAIR's research platform for object detection research, implementing popular algorithms like Mask R-CNN and RetinaNet.
-
9 Feb 2021 16 repositories listed Syntology ran 35 of 43 samples · 8 unverified · 14 pointer-only (licence)We present a convolution-free approach to video classification built exclusively on self-attention over space and time.
-
24 Jun 2021 15 repositories listed Syntology ran 7 of 32 samples · 25 unverifiedThe vision community is witnessing a modeling shift from CNNs to Transformers, where pure Transformer architectures have attained top accuracy on the major video recognition benchmarks.
-
8 May 2017 11 repositories listedFurthermore, based on the temporal segment networks, we won the video classification track at the ActivityNet challenge 2016 among 24 teams, which demonstrates the effectiveness of TSN and the proposed good practices.
-
19 Nov 2015 11 repositories listedOne of the challenges in modeling cognitive events from electroencephalogram (EEG) data is finding representations that are invariant to inter- and intra-subject differences, as well as to inherent noise associated with…
-
29 Mar 2021 10 repositories listed Syntology ran 14 of 21 samples · 7 unverified · 1 pointer-only (licence)We present pure-transformer based models for video classification, drawing upon the recent success of such models in image classification.
-
10 Apr 2020 10 repositories listed Syntology ran 2 of 22 samples · 20 unverifiedTherefore, in the present paper, we conduct exploration study in order to improve spatiotemporal 3D CNNs as follows: (i) Recently proposed large-scale video datasets help improve spatiotemporal 3D CNNs in terms of video…
-
2 Dec 2021 9 repositories listedIn this paper, we study Multiscale Vision Transformers (MViTv2) as a unified architecture for image and video classification, as well as object detection.
-
24 Jan 2022 8 repositories listedDifferent from the typical transformer blocks, the relation aggregators in our UniFormer block are equipped with local and global token affinity respectively in shallow and deep layers, allowing to tackle both…
-
9 Apr 2020 8 repositories listed Syntology ran 2 of 15 samples · 13 unverifiedThis paper presents X3D, a family of efficient video networks that progressively expand a tiny 2D image classification architecture along multiple network axes, in space, time, width and depth.
-
4 Apr 2019 7 repositories listed Syntology ran 1 of 4 samples · 3 unverifiedIt is natural to ask: 1) if group convolution can help to alleviate the high computational cost of video classification networks; 2) what factors matter the most in 3D group convolutional networks; and 3) what are good…
-
27 Sep 2016 7 repositories listed Syntology ran 1 of 9 samples · 8 unverifiedDespite the size of the dataset, some of our models train to convergence in less than a day on a single machine using TensorFlow.
-
9 Jun 2014 7 repositories listed Syntology ran 1 of 7 samples · 6 unverified · 2 pointer-only (licence)Our architecture is trained and evaluated on the standard video actions benchmarks of UCF-101 and HMDB-51, where it is competitive with the state of the art.
-
24 Apr 2018 6 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedIn this paper, we introduce a network architecture that takes long-term content into account and enables fast per-video processing at the same time.
-
4 Jul 2022 5 repositories listedIn this study, we focus on transferring knowledge for video classification tasks.
-
2 Oct 2018 5 repositories listedOur representation flow layer is a fully-differentiable layer designed to capture the `flow' of any representation channel within a convolutional neural network for action recognition.
-
27 Nov 2017 5 repositories listedIn this paper, however, we show that temporal information, especially longer-term patterns, may not be necessary to achieve competitive results on common video classification datasets.
-
21 Jun 2017 5 repositories listedIn particular, we evaluate our method on the large-scale multi-modal Youtube-8M v2 dataset and outperform all other methods in the Youtube 8M Large-Scale Video Understanding challenge.
-
9 Feb 2023 4 repositories listed Syntology ran 12 of 29 samples · 17 unverified · 26 pointer-only (licence)Reversible Vision Transformers achieve a reduced memory footprint of up to 15.
-
2 May 2019 4 repositories listedThis paper presents a study of semi-supervised learning with large convolutional networks.
-
30 Mar 2017 4 repositories listedWe demonstrate that using both RNNs (using LSTMs) and Temporal-ConvNets on spatiotemporal feature matrices are able to exploit spatiotemporal dynamics to improve the overall performance.
-
12 Jul 2023 3 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 2 pointer-only (licence)The ubiquitous and demonstrably suboptimal choice of resizing images to a fixed resolution before processing them with computer vision models has not yet been successfully challenged.
-
5 Aug 2021 3 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)It is worth noticing that our TokShift transformer is a pure convolutional-free video transformer pilot with computational efficiency for video understanding.
-
13 Mar 2021 3 repositories listedUsing improved training and scaling strategies, we design a family of ResNet architectures, ResNet-RS, which are 1.
-
7 Dec 2020 3 repositories listed Syntology ran 2 of 3 samples · 1 unverifiedRecent data augmentation strategies have been reported to address the overfitting problems in static image classifiers.
-
20 Jun 2020 3 repositories listedThis work introduces pyramidal convolution (PyConv), which is capable of processing the input at multiple filter scales.
-
1 Jun 2020 3 repositories listedThe core of our method is learnable and data-adaptive bilinear attentional transform (BA-Transform), whose merits are three-folds: first, BA-Transform is versatile to model a wide spectrum of local or global attentional…
-
16 Mar 2020 3 repositories listedIn this paper we challenge the common assumption that convolutional layers in modern CNNs are translation invariant.
-
2 Dec 2019 3 repositories listed Syntology ran 2 of 10 samples · 8 unverified · 3 pointer-only (licence)We empirically demonstrate a general and robust grid schedule that yields a significant out-of-the-box training speedup without a loss in accuracy for different models (I3D, non-local, SlowFast), datasets (Kinetics,…
Syntology lines on 16 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections