Browse State-of-the-Art › Action Classification
Action Classification
251 papers with code · 28 benchmarks · 33 datasets archive 2025-07-28
Image source: The Kinetics Human Action Video Dataset
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
28 leaderboard tables shown for this task, 28 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 28 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
33 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 33 until expanded.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 251 papers with code (457 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
31 Dec 2018 45 repositories listed Syntology ran 5 of 23 samples · 18 unverified · 3 pointer-only (licence)Accurate depth estimation from images is a fundamental task in many applications including scene understanding and reconstruction.
-
22 May 2017 34 repositories listed Syntology ran 16 of 28 samples · 12 unverified · 7 pointer-only (licence)The paucity of videos in current action classification datasets (UCF-101 and HMDB-51) has made it difficult to identify good video architectures, as most methods obtain similar performance on existing small-scale…
-
21 Nov 2017 32 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)Both convolutional and recurrent operations are building blocks that process one local neighborhood at a time.
-
Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks with Octave Convolution10 Apr 2019 28 repositories listed Syntology ran 14 of 34 samples · 20 unverified · 8 pointer-only (licence)Similarly, the output feature maps of a convolution layer can also be seen as a mixture of information at different frequencies.
-
30 Nov 2017 24 repositories listed Syntology ran 1 of 4 samples · 3 unverified · 4 pointer-only (licence)In this paper we discuss several forms of spatiotemporal convolutions for video analysis and study their effects on action recognition.
-
18 Nov 2021 23 repositories listed Syntology ran 3 of 30 samples · 27 unverifiedThree main techniques are proposed: 1) a residual-post-norm method combined with cosine attention to improve training stability; 2) A log-spaced continuous position bias method to effectively transfer models pre-trained…
-
2 Aug 2016 22 repositories listed Syntology ran 2 of 24 samples · 22 unverified · 3 pointer-only (licence)The other contribution is our study on a series of good practices in learning ConvNets on video data with the help of temporal segment network.
-
9 Feb 2021 16 repositories listed Syntology ran 35 of 43 samples · 8 unverified · 14 pointer-only (licence)We present a convolution-free approach to video classification built exclusively on self-attention over space and time.
-
24 Jun 2021 15 repositories listed Syntology ran 7 of 32 samples · 25 unverifiedThe vision community is witnessing a modeling shift from CNNs to Transformers, where pure Transformer architectures have attained top accuracy on the major video recognition benchmarks.
-
10 Dec 2018 15 repositories listed Syntology ran 0 of 10 samples · 10 unverifiedWe present SlowFast networks for video recognition.
-
20 Nov 2018 13 repositories listed Syntology ran 6 of 16 samples · 10 unverified · 4 pointer-only (licence)The explosive growth in video streaming gives rise to challenges on performing video understanding at high accuracy and low computation cost.
-
19 May 2017 13 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedWe describe the DeepMind Kinetics human action video dataset.
-
21 Jun 2021 11 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedIn this paper, we introduce a novel visual representation learning which relies on a handful of adaptively learned tokens, and which is applicable to both image and video understanding tasks.
-
8 May 2017 11 repositories listedFurthermore, based on the temporal segment networks, we won the video classification track at the ActivityNet challenge 2016 among 24 teams, which demonstrates the effectiveness of TSN and the proposed good practices.
-
29 Mar 2021 10 repositories listed Syntology ran 14 of 21 samples · 7 unverified · 1 pointer-only (licence)We present pure-transformer based models for video classification, drawing upon the recent success of such models in image classification.
-
23 Mar 2022 9 repositories listed Syntology ran 9 of 13 samples · 4 unverified · 12 pointer-only (licence)Pre-training video transformers on extra large-scale datasets is generally required to achieve premier performance on relatively small datasets.
-
2 Dec 2021 9 repositories listedIn this paper, we study Multiscale Vision Transformers (MViTv2) as a unified architecture for image and video classification, as well as object detection.
-
30 Nov 2018 9 repositories listed Syntology ran 8 of 15 samples · 7 unverified · 4 pointer-only (licence)In this work, we propose a new approach for reasoning globally in which a set of features are globally aggregated over the coordinate space and then projected to an interaction space where relational reasoning can be…
-
22 Apr 2021 8 repositories listed Syntology ran 13 of 26 samples · 13 unverified · 5 pointer-only (licence)We evaluate this fundamental architectural prior for modeling the dense nature of visual signals for a variety of video recognition tasks where it outperforms concurrent vision transformers that rely on large scale…
-
9 Apr 2020 8 repositories listed Syntology ran 2 of 15 samples · 13 unverifiedThis paper presents X3D, a family of efficient video networks that progressively expand a tiny 2D image classification architecture along multiple network axes, in space, time, width and depth.
-
4 Apr 2019 7 repositories listed Syntology ran 1 of 4 samples · 3 unverifiedIt is natural to ask: 1) if group convolution can help to alleviate the high computational cost of video classification networks; 2) what factors matter the most in 3D group convolutional networks; and 3) what are good…
-
9 Jun 2014 7 repositories listed Syntology ran 1 of 7 samples · 6 unverified · 2 pointer-only (licence)Our architecture is trained and evaluated on the standard video actions benchmarks of UCF-101 and HMDB-51, where it is competitive with the state of the art.
-
14 Nov 2022 6 repositories listed Syntology ran 1 of 3 samples · 2 unverifiedWe launch EVA, a vision-centric foundation model to explore the limits of visual representation at scale using only publicly accessible data.
-
4 May 2022 6 repositories listed Syntology ran 9 of 17 samples · 8 unverifiedWe apply a contrastive loss between unimodal image and text embeddings, in addition to a captioning loss on the multimodal decoder outputs which predicts text tokens autoregressively.
-
16 Dec 2021 6 repositories listedWe present Masked Feature Prediction (MaskFeat) for self-supervised pre-training of video models.
-
24 Apr 2018 6 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedIn this paper, we introduce a network architecture that takes long-term content into account and enables fast per-video processing at the same time.
-
31 Dec 2022 5 repositories listedIn this paper, we propose a novel framework called BIKE, which utilizes the cross-modal bridge to explore bidirectional knowledge: i) We introduce the Video Attribute Association mechanism, which leverages the…
-
4 Jul 2022 5 repositories listedIn this study, we focus on transferring knowledge for video classification tasks.
-
3 Sep 2021 5 repositories listedA recent work from Bello shows that training and scaling strategies may be more significant than model architectures for visual recognition.
-
22 Apr 2021 5 repositories listed Syntology ran 5 of 8 samples · 3 unverified · 8 pointer-only (licence)We train VATT end-to-end from scratch using multimodal contrastive losses and evaluate its performance by the downstream tasks of video action recognition, audio event classification, image classification, and…
Syntology lines on 24 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections