Home › Datasets › task › Video Classification
Video Classification datasets
archive 2025-07-28
22 datasets carry the task tag "Video Classification" (the task itself: Video Classification), ordered by the archive's paper count. Page 1 of 1: 22 shown of 22. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Video Classification datasets 1–22 of 22
Kinetics (Kinetics Human Action Video Dataset)
The Kinetics dataset is a large-scale, high-quality dataset for human action recognition in videos.
1,341 papers · 18 benchmarks
The Charades dataset is composed of 9,848 videos of daily indoors activities with an average length of 30 seconds, involving interactions with 46 objects classes in 15 types of indoor scenes and containing a vocabulary of 30 verbs leading…
428 papers · 6 benchmarks
The 20BN-SOMETHING-SOMETHING V2 dataset is a large collection of labeled video clips that show humans performing pre-defined basic actions with everyday objects.
290 papers · 7 benchmarks
The Breakfast Actions Dataset comprises of 10 actions related to breakfast preparation, performed by 52 different individuals in 18 different kitchens.
179 papers · 6 benchmarks
The YouTube-8M dataset is a large scale video dataset, which includes more than 7 million videos with 4716 classes labeled by the annotation system.
147 papers · 2 benchmarks
The 20BN-SOMETHING-SOMETHING dataset is a large collection of labeled video clips that show humans performing pre-defined basic actions with everyday objects.
117 papers · 3 benchmarks
The COIN dataset (a large-scale dataset for COmprehensive INstructional video analysis) consists of 11,827 videos related to 180 different tasks in 12 domains (e.g., vehicles, gadgets, etc.) related to our daily life.
105 papers · 2 benchmarks
Contains 68,536 activity instances in 68.8 hours of first and third-person video, making it one of the largest and most diverse egocentric datasets available.
33 papers · 1 benchmark
Home Action Genome is a large-scale multi-view video database of indoor daily activities.
9 papers · 2 benchmarks
MOMA-LRG (Multi-Object Multi-Actor activity parsing with Language-Refined Graphs)
A dataset dedicated to multi-object, multi-actor activity parsing.
3 papers · 1 benchmark
Whereas the action recognition community has focused mostly on detecting simple actions like clapping, walking or jogging, the detection of fights or in general aggressive behaviors has been comparatively less studied.
2 papers · 1 benchmark
MetaVD is a Meta Video Dataset for enhancing human action recognition datasets.
2 papers · 0 benchmarks
MoB (Malicious or Benign Cartoon Videos)
A dataset of cartoon video clips.
2 papers · 1 benchmark
This large collection of over 161,000 video-label pairs of video clips, shows humans drawing letters and digits in the air, and is used to evaluate a model’s ability to classify articulated motions correctly.
1 paper · 0 benchmarks
Bukva (Bukva: Russian Sign Language Alphabet)
We introduce a video dataset Bukva for Russian Dactyl Recognition task.
1 paper · 1 benchmark
Dataset for multimodal skills assessment focusing on assessing piano player’s skill level.
1 paper · 3 benchmarks
APPROVE consists of curated YouTube videos annotated with educational content.
1 paper · 1 benchmark
Trailers12k is a movie trailer dataset comprised of 12,000 titles associated to ten genres.
1 paper · 0 benchmarks
First of its kind paired win-fail action understanding dataset with samples from the following domains: “General Stunts,” “Internet Wins-Fails,” “Trick Shots,” & “Party Games.” The task is to identify successful and failed attempts at…
1 paper · 2 benchmarks
Crowd 11 (A Dataset for Fine Grained Crowd Behaviour Analysis)
This dataset defines a total of 11 crowd motion patterns and it is composed of over 6000 video sequences with an average length of 100 frames per sequence.
0 papers · 0 benchmarks
Laser powder bed fusion (LBPF) is the additive manufacturing (3D printing) process for metals.
0 papers · 0 benchmarks
Video Dataset (Storytelling Video Dataset (Russian, Emotion, Gesture, Speech))
The Storytelling Video Dataset is a high-quality, human-reviewed multimodal dataset featuring over 700 full-body video recordings of native Russian speakers.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.