Home › Datasets › modality › Videos
Videos datasets
archive 2025-07-28
1,014 datasets carry the modality tag "Videos", ordered by the archive's paper count. Page 18 of 22: 48 shown of 1,014. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Videos datasets 817–864 of 1,014
Characterising multimedia content with relevant, reliable and discriminating tags is vital for multimedia information retrieval.
1 paper · 0 benchmarks
Here we release the dataset (MultiChannelGrid, abbreviated as MCGrid) used in our paper LIMUSE: LIGHTWEIGHT MULTI-MODAL SPEAKER EXTRACTION](https://arxiv.org/abs/2111.04063)).
1 paper · 0 benchmarks
METEOR is a complex traffic dataset which captures traffic patterns in unstructured scenarios in India.
1 paper · 0 benchmarks
METU-VIREF is a video referring expression dataset comprising of videos from VIRAT Ground and ILSVRC2015 VID datasets.
1 paper · 0 benchmarks
MFA (Many Faces of Anger)
The MFA (Many Faces of Anger) dataset includes 200 in-the-wild videos from North American and Persian cultures with fine-grained labels of: 'annoyed', 'anger', 'disgust', 'hatred' and 'furious' and 13 related emojis.
1 paper · 1 benchmark
MH-FED (Meta Human Facial Expression Dataset)
This dataset provides a collection of 162K images and 70 Videos of Meta-Humans.
1 paper · 0 benchmarks
Consists of manually annotated dangerous and non-dangerous Kiki challenge videos.
1 paper · 0 benchmarks
Brazilian Sign Language (Libras) data set with 20 signs for sign language and gesture recognition benchmark: - Acontecer (To happen) - Aluno (Student) - Amarelo (Yellow) - América (America) - Aproveitar (To enjoy) - Bala (Candy) - Banco…
1 paper · 1 benchmark
MINT (a Multi-modal Image and Narrative Text Dubbing Dataset)
Foley audio, critical for enhancing the immersive experience in multimedia content, faces significant challenges in the AI-generated content (AIGC) landscape.
1 paper · 0 benchmarks
MIRACL-VC1 is a lip-reading dataset including both depth and color images.
1 paper · 1 benchmark
MLB (Mouse Lockbox Dataset)
This dataset provides high-resolution videos recorded from three perspectives with more than 110 hours of total playtime showing mice solving complex tasks.
1 paper · 0 benchmarks
MlGesture is a dataset for hand gesture recognition tasks, recorded in a car with 5 different sensor types at two different viewpoints.
1 paper · 0 benchmarks
MOD20 is an action recognition dataset consisting of videos collected from YouTube and our own drone.
1 paper · 0 benchmarks
MOET a dataset consists of gaze data from participants tracking specific objects, annotated with labels and bounding boxes, in crowded real-world videos, for training and evaluating attention decoding algorithms.
1 paper · 0 benchmarks
MOViD-A is a video-based synthesized dataset.
1 paper · 0 benchmarks
MSVD-Indonesian is derived from the MSVD dataset, which is obtained with the help of a machine translation service.
1 paper · 3 benchmarks
Contains about 1, 000 videos from 10 queries and their video tags, manual annotations, and associated web images.
1 paper · 0 benchmarks
MVX incorporates realistic physical world simulation with a differentiable accurate ray tracing wireless simulation that includes multi-agent and multimodal datasets for AI-driven digital twin applications in vehicular communication…
1 paper · 1 benchmark
The MedVidCL dataset contains a collection of 6, 617 videos annotated into ‘medical instructional’, ‘medical non-instructional' and ‘non-medical’ classes.
1 paper · 0 benchmarks
Mediapi-RGB is a bilingual corpus of French Sign Language (LSF) and written French in the form of subtitled videos, accompanied by complementary data (various representations, segmentation, vocabulary, etc.).
1 paper · 1 benchmark
Collecting data with a HIKVISION USB Camera DS-E11, we build a dataset called MentalHAD with four abnormal actions (climbing walls, hitting windows, climbing, and hitting) and six normal actions (crouching, standing, sitting, hand waving,…
1 paper · 0 benchmarks
Metaphorics is a newly introduced non-contextual skeleton action dataset.
1 paper · 0 benchmarks
MobiFace is the first dataset for single face tracking in mobile situations.
1 paper · 0 benchmarks
This dataset was generated to characterize mouse grooming behavior.
1 paper · 0 benchmarks
A large, annotated video dataset of mice performing a sequence of actions.
1 paper · 0 benchmarks
MuSoHu (Toward human-like social robot navigation: A large-scale, multi-modal, social human navigation dataset)
A large-scale, egocentric, multimodal, and context-aware dataset of human demonstrations of social navigation.
1 paper · 0 benchmarks
A dataset of music videos with continuous valence/arousal ratings as well as emotion tags.
1 paper · 0 benchmarks
A new multi-view egocentric dataset, Multi-Ego.
1 paper · 0 benchmarks
MultiSum is a dataset for multimodal summarization (MSMO).
1 paper · 0 benchmarks
Dataset for multimodal skills assessment focusing on assessing piano player’s skill level.
1 paper · 3 benchmarks
MuscleMap136 is a dataset for video-based Activated Muscle Group Estimation (AMGE) aiming at identifying currently activated muscular regions of humans performing a specific activity.
1 paper · 0 benchmarks
MuseChat Dataset (MuseChat: A Conversational Music Recommendation System for Videos (CVPR 2024 Highlight Paper))
Music recommendation for videos attracts growing interest in multi-modal research.
1 paper · 0 benchmarks
NES-VMDB is a dataset containing 98,940 gameplay videos from 389 NES games, each paired with its original soundtrack in symbolic format (MIDI).
1 paper · 0 benchmarks
In the last two years, millions of lives have been lost due to COVID-19.
1 paper · 0 benchmarks
Motion similarity annotations for NTU RGB+D 120 dataset to evaluate motion similarity in the real world.
1 paper · 0 benchmarks
This csv consists of (x-position, y-position, area) tuples of three views (left, middle, right) of downscaled binary masks with aspect ratio kept (64 x 128) from the 2019 YouTube-VIS challenge, which can be found at…
1 paper · 1 benchmark
Near-Collision is a large-scale dataset of 13,658 egocentric video snippets of humans navigating in indoor hallways.
1 paper · 0 benchmarks
Onchocerciasis is causing blindness in over half a million people in the world today.
1 paper · 0 benchmarks
NurViD (A Large Expert-Level Video Database for Nursing Procedure Activity Understanding)
We propose NurViD, a large video dataset with expert-level annotation for nursing procedure activity understanding.
1 paper · 0 benchmarks
OSTD (Open-Source-Total-Distance)
This dataset consists of 18 movies with duration range between 10 and 104 minutes leveraged from the OVSD dataset (Rotman et al., 2016).
1 paper · 0 benchmarks
🏃♂️ Open-HypermotionX Dataset Open-Hypermotion is a large-scale, high-quality dataset designed for training and evaluating pose-guided human image animation models, with a special focus on complex, dynamic human motions (Hypermotion),…
1 paper · 0 benchmarks
We create the first open-source large-scale S2V generation dataset OpenS2V-5M, which consists of five million high-quality 720P subject-text-video triples.
1 paper · 1 benchmark
The Oxford Town Center dataset is a 5-minute video with 7500 frames annotated, which is divided into 6500 for training and 1000 for testing data for pedestrian detection.
1 paper · 1 benchmark
The POTUS Corpus is a Database of Weekly Addresses for the Study of Stance in Politics and Virtual Agents.
1 paper · 0 benchmarks
PPED (Periodic Phenomena Event-based Dataset)
PPED: Periodic Phenomena Event-based Dataset The dataset features 12 one-second sequences of periodic phenomena (rotation - 01-06, flicker - 07-08, vibration - 09-10 and movement - 11-12) with GT frequencies ranging from 3.2Hz up to 2000Hz…
1 paper · 0 benchmarks
PTVD is a plot-oriented multimodal dataset in the TV domain.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.