Home › Datasets › task › Action Classification

Action Classification datasets

archive 2025-07-28

33 datasets carry the task tag "Action Classification" (the task itself: Action Classification), ordered by the archive's paper count. Page 1 of 1: 33 shown of 33. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Action Classification datasets 1–33 of 33

UCF101 (UCF101 Human Actions dataset)
UCF101 dataset is an extension of UCF50 and consists of 13,320 video clips, which are classified into 101 categories.
1,863 papers · 23 benchmarks
Kinetics (Kinetics Human Action Video Dataset)
The Kinetics dataset is a large-scale, high-quality dataset for human action recognition in videos.
1,341 papers · 18 benchmarks
The HMDB51 dataset is a large collection of realistic videos from various sources, including movies and web videos.
839 papers · 10 benchmarks
The ActivityNet dataset contains 200 different types of activities and a total of 849 hours of videos collected from YouTube.
807 papers · 17 benchmarks
The dataset contains 400 human action classes, with at least 400 video clips for each action.
712 papers · 0 benchmarks
The Charades dataset is composed of 9,848 videos of daily indoors activities with an average length of 30 seconds, involving interactions with 46 objects classes in 15 types of indoor scenes and containing a vocabulary of 30 verbs leading…
428 papers · 6 benchmarks
The THUMOS14 (THUMOS 2014) dataset is a large-scale video dataset that includes 1,010 videos for validation and 1,574 videos for testing from 20 classes.
318 papers · 18 benchmarks
The 20BN-SOMETHING-SOMETHING V2 dataset is a large collection of labeled video clips that show humans performing pre-defined basic actions with everyday objects.
290 papers · 7 benchmarks
YouCook2 is the largest task-oriented, instructional video dataset in the vision community.
198 papers · 7 benchmarks
The Kinetics-600 is a large-scale action recognition dataset which consists of around 480K videos from 600 action categories.
148 papers · 3 benchmarks
Kinetics-700 is a video dataset of 650,000 clips that covers 700 human action classes.
95 papers · 3 benchmarks
BABEL is a large dataset with language labels describing the actions being performed in mocap sequences.
72 papers · 1 benchmark
WLASL (Word-Level American Sign Language)
WLASL is a large video dataset for Word-Level American Sign Language (ASL) recognition, which features 2,000 common different words in ASL.
66 papers · 3 benchmarks
A novel large-scale corpus of manual annotations for the SoccerNet video dataset, along with open challenges to encourage more research in soccer understanding and broadcast production.
58 papers · 6 benchmarks
A benchmark for action spotting in soccer videos.
53 papers · 1 benchmark
CelebV-HQ is a large-scale video facial attributes dataset with annotations.
36 papers · 3 benchmarks
A large scale dataset with daily-living activities performed in a natural manner.
31 papers · 1 benchmark
Toyota Smarthome dataset (Toyota Smarthome Trimmed)
Toyota Smarthome Trimmed has been designed for the activity classification task of 31 activities.
23 papers · 0 benchmarks
Jester Gesture Recognition dataset includes 148,092 labeled video clips of humans performing basic, pre-defined hand gestures in front of a laptop camera or webcam.
16 papers · 6 benchmarks
A database with 2,000 videos captured by surveillance cameras in real-world scenes.
16 papers · 1 benchmark
HAA500 (Human-Centric Atomic Action Dataset)
HAA500 is a manually annotated human-centric atomic action dataset for action recognition on 500 classes with over 591k labeled frames.
14 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
11 papers · 1 benchmark
Is a collection of action videos from many different countries.
9 papers · 1 benchmark
We describe the 2020 edition of the DeepMind Kinetics human action dataset, which replenishes and extends the Kinetics-700 dataset.
9 papers · 1 benchmark
The Sims4Action Dataset: a videogame-based dataset for Synthetic→Real domain adaptation for human activity recognition.
5 papers · 0 benchmarks
Comprises of 171,191 video segments from 346 high-quality soccer games.
5 papers · 0 benchmarks
WiGesture (Wireless Sensing Dataset for Gesture Recognition and People ID Identification with ESP32)
WiGesture dataset contains data related to gesture recognition and people id identification in a meeting room scenario.
5 papers · 2 benchmarks
TTStroke-21 ME21 (TTStroke-21 for MediaEval 2021)
This task offers researchers an opportunity to test their fine-grained classification methods for detecting and recognizing strokes in table tennis videos.
3 papers · 2 benchmarks
TTStroke-21 ME22 (TTStroke-21 for MediaEval 2022)
TTStroke-21 for MediaEval 2022.
3 papers · 2 benchmarks
WiFall (Wireless Sensing Dataset for Fall Detection, Action Recognition and People ID Identification with ESP32-S3)
WiFall dataset contains data related to fall detection, action recognition and people id identification in a meeting room scenario.
2 papers · 1 benchmark
MLB (Mouse Lockbox Dataset)
This dataset provides high-resolution videos recorded from three perspectives with more than 110 hours of total playtime showing mice solving complex tasks.
1 paper · 0 benchmarks
First of its kind paired win-fail action understanding dataset with samples from the following domains: “General Stunts,” “Internet Wins-Fails,” “Trick Shots,” & “Party Games.” The task is to identify successful and failed attempts at…
1 paper · 2 benchmarks
InfiniteRep is a synthetic, open-source dataset for fitness and physical therapy (PT) applications.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.