Browse State-of-the-Art › Action Recognition

Action Recognition

1,058 papers with code · 56 benchmarks · 115 datasets archive 2025-07-28

Computer VisionTime Series

Action Recognition is a computer vision task that involves recognizing human actions in videos or images. The goal is to classify and categorize the actions being performed in the video or image into a predefined set of action classes.

In the video domain, it is an open question whether training an action classification network on a sufficiently large dataset, will give a similar boost in performance when applied to a different temporal task or dataset. The challenges of building video datasets has meant that most popular benchmarks for action recognition are small, having on the order of 10k videos.

Please note some benchmarks may be located in the Action Classification or Video Classification tasks, e.g. Kinetics-400.

Description from the archive archive 2025-07-28; Papers-with-Code links inside it are rewritten to this site.

Benchmarks archive 2025-07-28

56 leaderboard tables shown for this task, 56 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 56 until expanded.

DatasetBest model (first row in archive order)PaperCodeSyntologyCompare
Something-Something V2 (123 rows) MVD (Kinetics400 pretrain, ViT-H, 16 frame) Masked Video Distillation: Rethinking Masked Feature Modeling for... code — Compare
UCF101 (91 rows) FTP-UniFormerV2-L/14 Enhancing Video Transformers for Action Understanding with... — — Compare
HMDB-51 (77 rows) VideoMAE V2-g VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking code Syntology ran 2 of 6 samples · 4 unverified Compare
Something-Something V1 (74 rows) InternVideo InternVideo: General Video Foundation Models via Generative and... code Syntology ran 3 of 3 samples · 0 unverified Compare
AVA v2.2 (38 rows) LART (Hiera-H, K700 PT+FT) On the Benefits of 3D Pose and Tracking for Human Action Recognition code — Compare
EPIC-KITCHENS-100 (32 rows) LLaVAction LLaVAction: evaluating and training multi-modal large language... code — Compare
NTU RGB+D (28 rows) DSCNet (RGB + Pose) A Dense-Sparse Complementary Network for Human Action Recognition... code — Compare
NTU RGB+D 120 (21 rows) DSCNet (RGB + Pose) A Dense-Sparse Complementary Network for Human Action Recognition... code — Compare
Diving-48 (18 rows) LVMAE Extending Video Masked Autoencoders to 128 frames — — Compare
ActivityNet (16 rows) Text4Vis (w/ ViT-L) Revisiting Classifier: Transferring Vision-Language Models for... code — Compare
AVA v2.1 (15 rows) STAR/L End-to-End Spatio-Temporal Action Localisation with Video Transformers — — Compare
H2O (2 Hands and Objects) (11 rows) HandFormer-B/21x8 On the Utility of 3D Hand Poses for Action Recognition code Syntology ran 8 of 13 samples · 5 unverified Compare
THUMOS’14 (10 rows) BMN BMN: Boundary-Matching Network for Temporal Action Proposal Generation code Syntology ran 3 of 11 samples · 8 unverified Compare
Sports-1M (9 rows) ip-CSN-152 (RGB) Video Classification with Channel-Separated Convolutional Networks code Syntology ran 1 of 4 samples · 3 unverified Compare
HACS (8 rows) InternVideo2-6B InternVideo2: Scaling Foundation Models for Multimodal Video Understanding code — Compare
Charades-Ego (6 rows) LaViLa (Finetuned, TimeSformer-L) Learning Video Representations from Large Language Models code Syntology ran 6 of 20 samples · 14 unverified Compare
Volleyball (4 rows) PoseC3D (Pose Only) Revisiting Skeleton-based Action Recognition code — Compare
Animal Kingdom (4 rows) ARTEMIS: animal recognition through enhanced multimodal integration system — — — Compare
BAR (4 rows) DebiAN Discover and Mitigate Unknown Biases with Debiasing Alternate Networks code Syntology ran 2 of 9 samples · 7 unverified Compare
HAA500 (4 rows) TSN HAA500: Human-Centric Atomic Action Dataset with Curated Videos — — Compare
LoTE-Animal (4 rows) SlowOnly r50 LoTE-Animal: A Long Time-span Dataset for Endangered Animal... — — Compare
UAV-Human (4 rows) PMI Sampler PMI Sampler: Patch Similarity Guided Frame Selection for Aerial... code — Compare
Jester (Gesture Recognition) (3 rows) DirecFormer DirecFormer: A Directed Attention in Transformer Approach to... code Syntology ran 4 of 8 samples · 4 unverified Compare
RareAct (3 rows) 🦩 Flamingo Flamingo: a Visual Language Model for Few-Shot Learning code Syntology ran 18 of 24 samples · 6 unverified Compare
Real Life Violence Situations Dataset (3 rows) DeVTr Data Efficient Video Transformer for Violence Detection code — Compare
ICVL-4 (2 rows) OHA-GCN (Two stream; HP + OHP-hands + informative samples) Skeleton-based Action Recognition of People Handling Objects — — Compare
IRD (2 rows) OHA-GCN (Two stream; HP + OHP-hands + informative samples) Skeleton-based Action Recognition of People Handling Objects — — Compare
miniSports (2 rows) IF+MD+RGB-R (ResNet-18) SCSampler: Sampling Salient Clips from Video for Efficient Action... — — Compare
UCF-101 (2 rows) DMC-Net (ResNet-18) DMC-Net: Generating Discriminative Motion Cues for Fast Compressed... — — Compare
Drone-Action (2 rows) FAR FAR: Fourier Aerial Video Recognition code — Compare
Mimetics (2 rows) JMRN Pose And Joint-Aware Action Recognition code — Compare
Okutama-Action (2 rows) PLAR with bbox (Ours) SCP: Soft Conditional Prompt Learning for Aerial Video Action Recognition — — Compare
Penn Action (2 rows) 3DA (RGB + Pose) Cross-Modal Learning with 3D Deformable Attention for Action Recognition — — Compare
SL-Animals (2 rows) SEW-Resnet18 (3sets) EventRPG: Event Data Augmentation with Relevance Propagation Guidance code Syntology ran 5 of 16 samples · 11 unverified Compare
ActionNet-VE (1 row) Baseline Extensible Hierarchical Method of Detecting Interactive Actions... — — Compare
Charades (1 row) MSQNet Actor-agnostic Multi-label Action Recognition with Multi-modal Query code — Compare
EgoGesture (1 row) TSM+W3 Knowing What, Where and When to Look: Efficient Video Action... — — Compare
EPIC-KITCHENS-55 (1 row) TSM+W3 - full res Knowing What, Where and When to Look: Efficient Video Action... — — Compare
HMDB51 (1 row) MSQNet Actor-agnostic Multi-label Action Recognition with Multi-modal Query code — Compare
UTD-MHAD (1 row) Action Machine (RGB only) Action Machine: Rethinking Action Recognition in Trimmed Videos — — Compare
VIRAT Ground 2.0 (1 row) DHCM A Hierarchical Context Model for Event Recognition in Surveillance Video — — Compare
DVS128 Gesture (1 row) SEW-Resnet18 EventRPG: Event Data Augmentation with Relevance Propagation Guidance code Syntology ran 5 of 16 samples · 11 unverified Compare
Hockey (1 row) MSQNet Actor-agnostic Multi-label Action Recognition with Multi-modal Query code — Compare
IndustReal (1 row) MViT-V2 IndustReal: A Dataset for Procedure Step Recognition Handling... code Syntology ran 4 of 5 samples · 1 unverified Compare
KTH (1 row) CNN-GRU Temporal Relations of Informative Frames in Action Recognition code — Compare
MECCANO (1 row) SlowFast The MECCANO Dataset: Understanding Human-Object Interactions from... code — Compare
MTL-AQA (1 row) C3D-AVG What and How Well You Performed? A Multitask Learning Approach to... code — Compare
N-UCLA (1 row) DVANet DVANet: Disentangling View and Action Features for Multi-View... code — Compare
NEC Drone (1 row) FAR FAR: Fourier Aerial Video Recognition code — Compare
RoCoG-v2 (1 row) AZTR (Ours) AZTR: Aerial Video Action Recognition with Auto Zoom and Temporal Reasoning — — Compare
Skeleton-Mimetics (1 row) Structured Keypoint Pooling Unified Keypoint-based Action Recognition Framework via Structured... — — Compare
THUMOS14 (1 row) MSQNet Actor-agnostic Multi-label Action Recognition with Multi-modal Query code — Compare
UAV Human (1 row) FAR FAR: Fourier Aerial Video Recognition code — Compare
UCF 101 (1 row) R2+1D-BERT Late Temporal Modeling in 3D CNN Architectures with BERT for... code — Compare
UCFSports (1 row) CNN-LSTM Temporal Relations of Informative Frames in Action Recognition code — Compare
Win-Fail Action Understanding (1 row) 2DCNN+TRN Win-Fail Action Recognition code — Compare

Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.

Libraries

Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.

Datasets archive 2025-07-28

115 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 115 until expanded.

Subtasks archive 2025-07-28

15 subtasks in the archive's task tree.

Parent tasks archive 2025-07-28

Most implemented papers archive 2025-07-28

30 shown of 1,058 papers with code (2,759 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.

  • 28 May 2019 144 repositories listed Syntology ran 171 of 302 samples · 131 unverified · 112 pointer-only (licence)
    Convolutional Neural Networks (ConvNets) are commonly developed at a fixed resource budget, and then scaled up for better accuracy if more resources are available.
  • 26 Feb 2021 82 repositories listed Syntology ran 16 of 20 samples · 4 unverified · 16 pointer-only (licence)
    State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories.
  • 22 May 2017 34 repositories listed Syntology ran 16 of 28 samples · 12 unverified · 7 pointer-only (licence)
    The paucity of videos in current action classification datasets (UCF-101 and HMDB-51) has made it difficult to identify good video architectures, as most methods obtain similar performance on existing small-scale…
  • 21 Nov 2017 32 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)
    Both convolutional and recurrent operations are building blocks that process one local neighborhood at a time.
  • 2 Dec 2014 29 repositories listed
    We propose a simple, yet effective approach for spatiotemporal feature learning using deep 3-dimensional convolutional networks (3D ConvNets) trained on a large scale supervised video dataset.
  • 27 Nov 2017 26 repositories listed Syntology ran 7 of 8 samples · 1 unverified · 3 pointer-only (licence)
    The purpose of this study is to determine whether current video datasets have sufficient data for training very deep convolutional neural networks (CNNs) with spatio-temporal three-dimensional (3D) kernels.
  • 23 Jan 2018 24 repositories listed Syntology ran 2 of 2 samples · 0 unverified
    Dynamics of human body skeletons convey significant information for human action recognition.
  • 30 Nov 2017 24 repositories listed Syntology ran 1 of 4 samples · 3 unverified · 4 pointer-only (licence)
    In this paper we discuss several forms of spatiotemporal convolutions for video analysis and study their effects on action recognition.
  • 30 Oct 2017 24 repositories listed Syntology ran 4 of 7 samples · 3 unverified · 3 pointer-only (licence)
    Over the last decade, Convolutional Neural Network (CNN) models have been highly successful in solving complex vision problems.
  • 20 Nov 2014 24 repositories listed Syntology ran 12 of 32 samples · 20 unverified · 27 pointer-only (licence)
    We propose a novel paradigm for evaluating image descriptions that uses human consensus.
  • 2 Aug 2016 22 repositories listed Syntology ran 2 of 24 samples · 22 unverified · 3 pointer-only (licence)
    The other contribution is our study on a series of good practices in learning ConvNets on video data with the help of temporal segment network.
  • 9 Feb 2021 16 repositories listed Syntology ran 35 of 43 samples · 8 unverified · 14 pointer-only (licence)
    We present a convolution-free approach to video classification built exclusively on self-attention over space and time.
  • 24 Jun 2021 15 repositories listed Syntology ran 7 of 32 samples · 25 unverified
    The vision community is witnessing a modeling shift from CNNs to Transformers, where pure Transformer architectures have attained top accuracy on the major video recognition benchmarks.
  • 23 Jul 2019 15 repositories listed Syntology ran 3 of 11 samples · 8 unverified
    To address these difficulties, we introduce the Boundary-Matching (BM) mechanism to evaluate confidence scores of densely distributed proposals, which denote a proposal as a matching pair of starting and ending…
  • 10 Dec 2018 15 repositories listed Syntology ran 0 of 10 samples · 10 unverified
    We present SlowFast networks for video recognition.
  • 20 Nov 2018 13 repositories listed Syntology ran 6 of 16 samples · 10 unverified · 4 pointer-only (licence)
    The explosive growth in video streaming gives rise to challenges on performing video understanding at high accuracy and low computation cost.
  • 16 Feb 2015 12 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 2 pointer-only (licence)
    We further evaluate the representations by finetuning them for a supervised learning problem - human action recognition on the UCF-101 and HMDB-51 datasets.
  • 8 May 2017 11 repositories listed
    Furthermore, based on the temporal segment networks, we won the video classification track at the ActivityNet challenge 2016 among 24 teams, which demonstrates the effectiveness of TSN and the proposed good practices.
  • 29 Mar 2021 10 repositories listed Syntology ran 14 of 21 samples · 7 unverified · 1 pointer-only (licence)
    We present pure-transformer based models for video classification, drawing upon the recent success of such models in image classification.
  • 23 Mar 2022 9 repositories listed Syntology ran 9 of 13 samples · 4 unverified · 12 pointer-only (licence)
    Pre-training video transformers on extra large-scale datasets is generally required to achieve premier performance on relatively small datasets.
  • 2 Dec 2021 9 repositories listed
    In this paper, we study Multiscale Vision Transformers (MViTv2) as a unified architecture for image and video classification, as well as object detection.
  • 30 Nov 2018 9 repositories listed Syntology ran 8 of 15 samples · 7 unverified · 4 pointer-only (licence)
    In this work, we propose a new approach for reasoning globally in which a set of features are globally aggregated over the coordinate space and then projected to an interaction space where relational reasoning can be…
  • 23 May 2017 9 repositories listed
    The AVA dataset densely annotates 80 atomic visual actions in 430 15-minute video clips, where actions are localized in space and time, resulting in 1.
  • 22 Apr 2021 8 repositories listed Syntology ran 13 of 26 samples · 13 unverified · 5 pointer-only (licence)
    We evaluate this fundamental architectural prior for modeling the dense nature of visual signals for a variety of video recognition tasks where it outperforms concurrent vision transformers that rely on large scale…
  • 2 Jan 2021 8 repositories listed
    Convolutional Neural Networks (CNNs) use pooling to decrease the size of activation maps.
  • 5 Nov 2018 8 repositories listed
    In this paper, in contrast to the existing CNN+RNN or pure 3D convolution based approaches, we explore a novel spatial temporal network (StNet) architecture for both local and global spatial-temporal modeling in videos.
  • 23 Jun 2020 7 repositories listed Syntology ran 3 of 18 samples · 15 unverified
    This paper introduces the pipeline to extend the largest dataset in egocentric vision, EPIC-KITCHENS.
  • 4 Apr 2019 7 repositories listed Syntology ran 1 of 4 samples · 3 unverified
    It is natural to ask: 1) if group convolution can help to alleviate the high computational cost of video classification networks; 2) what factors matter the most in 3D group convolutional networks; and 3) what are good…
  • 14 Jan 2018 7 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)
    Over the past decade, multivariate time series classification has received great attention.
  • 27 Sep 2016 7 repositories listed Syntology ran 1 of 9 samples · 8 unverified
    Despite the size of the dataset, some of our models train to convergence in less than a day on a single machine using TensorFlow.

Syntology lines on 24 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections