Home › Datasets › modality › Actions
Actions datasets
archive 2025-07-28
67 datasets carry the modality tag "Actions", ordered by the archive's paper count. Page 1 of 2: 48 shown of 67. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Actions datasets 1–48 of 67
The Breakfast Actions Dataset comprises of 10 actions related to breakfast preparation, performed by 52 different individuals in 18 different kitchens.
179 papers · 6 benchmarks
NExT-QA is a VideoQA benchmark targeting the explanation of video contents.
174 papers · 3 benchmarks
BEAT (Body-Expression-Audio-Text)
BEAT has i) 76 hours, high-quality, multi-modal data captured from 30 speakers talking with eight different emotions and in four different languages, ii) 32 millions frame-level emotion and semantic relevance annotations.
56 papers · 1 benchmark
The KIT Motion-Language is a dataset linking human motion and natural language.
48 papers · 2 benchmarks
SBU-Kinect-Interaction dataset version 2.0 comprises of RGB-D video sequences of humans performing interaction activities that are recording using the Microsoft Kinect sensor.
28 papers · 4 benchmarks
100 tasks from LIBERO-100 suite.
22 papers · 1 benchmark
SCAND (Socially CompliAnt Navigation Dataset)
Have you wondered how autonomous mobile robots should share space with humans in public spaces?
18 papers · 0 benchmarks
V-D4RL provides pixel-based analogues of the popular D4RL benchmarking tasks, derived from the dmcontrol suite, along with natural extensions of two state-of-the-art online pixel-based continuous control algorithms, DrQ-v2 and DreamerV2,…
16 papers · 0 benchmarks
100 tasks from LIBERO-100 suite.
12 papers · 1 benchmark
UI-PRMD (University of Idaho – Physical Rehabilitation Movement Dataset)
UI-PRMD is a data set of movements related to common exercises performed by patients in physical therapy and rehabilitation programs.
10 papers · 2 benchmarks
Atari-HEAD is a dataset of human actions and eye movements recorded while playing Atari videos games.
9 papers · 0 benchmarks
WebLINX (Real-World Website Navigation with Multi-Turn)
WebLINX is a large-scale benchmark of 100K interactions across 2300 expert demonstrations of conversational web navigation.
6 papers · 1 benchmark
BRACE (The Breakdancing Competition Dataset for Dance Motion Synthesis)
BRACE is a dataset for audio-conditioned dance motion synthesis challenging common assumptions for this task: - strong music-dance correlation - controlled motion data - simple poses and movements To address these issues: - We focus on…
5 papers · 2 benchmarks
The Sims4Action Dataset: a videogame-based dataset for Synthetic→Real domain adaptation for human activity recognition.
5 papers · 0 benchmarks
The dataset uses VGG-Sound which consists of 10s clips collected from YouTube for 309 sound classes.
5 papers · 0 benchmarks
DoMSEV (Dataset of Multimodal Semantic Egocentric Video)
The Dataset of Multimodal Semantic Egocentric Video (DoMSEV) contains 80-hours of multimodal (RGB-D, IMU, and GPS) data related to First-Person Videos with annotations for recorder profile, frame scene, activities, interaction, and…
4 papers · 0 benchmarks
HowTo100M Adverbs is a subset from HowTo100M with mined adverbs from 83 tasks in HowTo100M.
4 papers · 1 benchmark
SportsPose (SportsPose - A Dynamic 3D sports pose dataset)
Accurate 3D human pose estimation is essential for sports analytics, coaching, and injury prevention.
4 papers · 0 benchmarks
The eSports Sensors dataset contains sensor data collected from 10 players in 22 matches in League of Legends.
4 papers · 2 benchmarks
ActivityNet Adverbs is a subset from the ActivityNet dataset with extracted verb-adverb annotations.
3 papers · 2 benchmarks
MSR-VTT Adverbs is a subset from MSR-VTT with extracted verb-adverb annotations.
3 papers · 2 benchmarks
This dataset contains a large set (~3.2 Million) of high quality expert trajectories generated from a geometrically consist hybrid planner in a wide variety of environment (~575,000 environments).
3 papers · 0 benchmarks
OREBA (Objectively Recognizing Eating Behavior and Associated Intake)
The OREBA dataset aims to provide a comprehensive multi-sensor recording of communal intake occasions for researchers interested in automatic detection of intake gestures.
3 papers · 0 benchmarks
SDN (Situated Dialogue Navigation)
Situated Dialogue Navigation (SDN) is a navigation benchmark of 183 trials with a total of 8415 utterances, around 18.7 hours of control streams, and 2.9 hours of trimmed audio.
3 papers · 0 benchmarks
UESTC-MMEA-CL (A multi-modal egocentric activity dataset for continual learning)
UESTC-MMEA-CL is a new multi-modal activity dataset for continual egocentric activity recognition, which is proposed to promote future studies on continual learning for first-person activity recognition in wearable applications.
3 papers · 0 benchmarks
VATEX Adverbs is a subset from VATEX with extracted verb-adverb annotations.
3 papers · 2 benchmarks
This package provides utilities for generation, filtering, solving, visualizing, and processing of mazes for training ML systems.
3 papers · 0 benchmarks
3DYoga90 (3DYoga90: A Hierarchical Video Dataset for Yoga Pose Understanding)
3DYoga90 is organized within a three-level label hierarchy.
2 papers · 0 benchmarks
CVB (Video Dataset of Cattle Visual Behaviors)
Existing image/video datasets for cattle behavior recognition are mostly small, lack well-defined labels, or are collected in unrealistic controlled environments.
2 papers · 0 benchmarks
MiniWob++ is a suite of web-browser based tasks introduced in Liu et al.
2 papers · 0 benchmarks
RHM (Rhm: Robot house multi-view human activity recognition dataset)
The Robot House Multi-View dataset (RHM) contains four views: Front, Back, Ceiling, and Robot Views.
2 papers · 1 benchmark
RL Unplugged is suite of benchmarks for offline reinforcement learning.
2 papers · 0 benchmarks
StarData is a StarCraft: Brood War replay dataset, with 65,646 games.
2 papers · 0 benchmarks
TI1K Dataset (Thumb Index 1000 Hand & Fingertip Detection Dataset)
Thumb Index 1000 (TI1K) is a dataset of 1000 hand images with the hand bounding box, and thumb and index fingertip positions.
2 papers · 0 benchmarks
In this dataset two robots, Baxter and UR5, perform 8 behaviors (look, grasp, pick, hold, shake, lower, drop, and push) on 95 objects that vary by 5 color (blue, green, red, white, and yellow), 6 contents (wooden button, plastic dices,…
1 paper · 0 benchmarks
This dataset contains the bus trajectory dataset collected by 6 volunteers who were asked to travel across the sub-urban city of Durgapur, India, on intra-city buses (route name: 54 Feet).
1 paper · 0 benchmarks
We present a new simulated dataset for pedestrian action anticipation collected using the CARLA simulator.
1 paper · 0 benchmarks
In this dataset an uppertorso humanoid robot with 7-DOF arm explored 100 different objects belonging to 20 different categories using 10 behaviors: Look, Crush, Grasp, Hold, Lift, Drop, Poke, Push, Shake and Tap.
1 paper · 0 benchmarks
This dataset contains Axivity AX3 wrist-worn activity tracker data that were collected from 151 participants in 2014-2016 around the Oxfordshire area.
1 paper · 0 benchmarks
Two versions of the dataset are offered: one is the full dataset used to train the models in DeformPAM, and the other is a mini dataset for easier examination.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
HA-ViD (HA-ViD: A Human Assembly Video Dataset)
Understanding comprehensive assembly knowledge from videos is critical for futuristic ultra-intelligent industry.
1 paper · 0 benchmarks
LARa (Logistic Activity Recognition Challenge)
LARa is the first freely accessible logistics-dataset for human activity recognition.
1 paper · 0 benchmarks
The dataset was collected from two courses offered on the University of Jordan's E-learning Portal during the second semester of 2020, namely "Computer Skills for Humanities Students" (CSHS) and "Computer Skills for Medical Students"…
1 paper · 0 benchmarks
LSDBench (Long-video Sampling Dilemma Benchmark)
A benchmark that focuses on the sampling dilemma in long-video tasks.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Whole-body, low-level control/manipulation demonstration dataset for ManiSkill-HAB.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.