Home › Datasets › modality › Videos
Videos datasets
archive 2025-07-28
1,014 datasets carry the modality tag "Videos", ordered by the archive's paper count. Page 12 of 22: 48 shown of 1,014. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Videos datasets 529–576 of 1,014
Large multimodal models (LMMs) are processing increasingly longer and richer inputs.
5 papers · 1 benchmark
iPer is a new dataset, with diverse styles of clothes in videos, for the evaluation of human motion imitation, appearance transfer, and novel view synthesis.
5 papers · 0 benchmarks
ARAUS (Affective Responses to Augmented Urban Soundscapes)
Choosing optimal maskers for existing soundscapes to effect a desired perceptual change via soundscape augmentation is non-trivial due to extensive varieties of maskers and a dearth of benchmark datasets with which to compare and develop…
4 papers · 0 benchmarks
ARMBench is a large-scale, object-centric benchmark dataset for robotic manipulation in the context of a warehouse.
4 papers · 1 benchmark
AnimeRun is a 2D animation visual correspondence dataset.
4 papers · 0 benchmarks
CCv2 (Casual Conversations v2)
Casual Conversations v2 (CCv2) is composed of over 5,567 participants (26,467 videos) and intended mainly to be used for assessing the performance of already trained models in computer vision and audio applications for the purposes…
4 papers · 0 benchmarks
The nine (moving camera) videos in this benchmark exhibit camouflaged animals that are difficult to see in a single frame, but can be detected based upon their motion across frames.
4 papers · 1 benchmark
The TUB CrowdFlow is a synthetic dataset that contains 10 sequences showing 5 scenes.
4 papers · 0 benchmarks
DoMSEV (Dataset of Multimodal Semantic Egocentric Video)
The Dataset of Multimodal Semantic Egocentric Video (DoMSEV) contains 80-hours of multimodal (RGB-D, IMU, and GPS) data related to First-Person Videos with annotations for recorder profile, frame scene, activities, interaction, and…
4 papers · 0 benchmarks
The images in DukeMTMC-attribute dataset comes from Duke University.
4 papers · 1 benchmark
The dataset contains 7000 videos: native, altered and exchanged through social platforms.
4 papers · 0 benchmarks
FR-FS (Fall Recognition in Figure Skating)
The FR-FS dataset contains 417 videos collected from FIV dataset and Pingchang 2018 Winter Olympic Games.
4 papers · 0 benchmarks
Goal is a novel dataset of football (or 'soccer') highlights videos with transcribed live commentaries in English.
4 papers · 0 benchmarks
Hate speech has become one of the most significant issues in modern society, with implications in both the online and offline worlds.
4 papers · 1 benchmark
HowTo100M Adverbs is a subset from HowTo100M with mined adverbs from 83 tasks in HowTo100M.
4 papers · 1 benchmark
HuPR (Human Pose with Millimeter Wave Radar)
HuPR is a human pose estimation benchmark is created using cross-calibrated mmWave radar sensors and a monocular RGB camera for cross-modality training of radar-based human pose estimation.
4 papers · 0 benchmarks
Human-Animal-Cartoon (HAC) dataset consists of seven actions (‘sleeping’, ‘watching tv’, ‘eating’, ‘drinking’, ‘swimming’, ‘running’, and ‘opening door’) performed by humans, animals, and cartoon figures, forming three different domains.
4 papers · 0 benchmarks
Kitchen Scenes is a multi-view RGB-D dataset of nine kitchen scenes, each containing several objects in realistic cluttered environments including a subset of objects from the BigBird dataset.
4 papers · 0 benchmarks
M³-VOS (M³-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation)
💡 Description A new benchmark, Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation (M³-VOS), to verify the ability of models to understand object phases, which consists of 479 high-resolution videos spanning over 10…
4 papers · 1 benchmark
The M5Product dataset is a large-scale multi-modal pre-training dataset with coarse and fine-grained annotations for E-products.
4 papers · 0 benchmarks
MA-52 (Micro-Action 52 dataset)
The MA-52 dataset provides the whole-body perspective including gestures, upper- and lower-limb movements, attempting to reveal comprehensive micro-action cues.
4 papers · 1 benchmark
MEVID (Multi-view Extended Videos with Identities Dataset)
Multi-view Extended Videos with Identities dataset (MEVID) is a dataset for large-scale, video person re-identification (ReID) in the wild.
4 papers · 0 benchmarks
MMToM-QA (Multimodal Theory of Mind Question Answering)
MMToM-QA is the first multimodal benchmark to evaluate machine Theory of Mind (ToM), the ability to understand people's minds.
4 papers · 0 benchmarks
MSRVTT-CTN Dataset This dataset contains CTN annotations for the MSRVTT-CTN benchmark dataset in JSON format.
4 papers · 1 benchmark
This is a dataset for video deinterlacing problem.
4 papers · 1 benchmark
MSVD-CTN (MSVD Causal-Temporal Narrative)
MSVD-CTN Dataset This dataset contains CTN annotations for the MSVD-CTN benchmark dataset in JSON format.
4 papers · 1 benchmark
MovingFashion is a dataset for video-to-shop, the task of retrieving clothes which are worn in social media videos.
4 papers · 1 benchmark
For each dataset we provide a short description as well as some characterization metrics.
4 papers · 0 benchmarks
PDEBench provides a diverse and comprehensive set of benchmarks for scientific machine learning, including challenging and realistic physical problems.
4 papers · 0 benchmarks
PHD² (Personalized Highlight Detection Dataset)
The dataset contains information on what video segments a specific user considers a highlight.
4 papers · 0 benchmarks
QUVA Repetition dataset consists of 100 videos displaying a wide variety of repetitive video dynamics, including swimming, stirring, cutting, combing and music-making.
4 papers · 0 benchmarks
Collects dense per-video-shot concept annotations.
4 papers · 1 benchmark
Text and signs around roads provide crucial information for drivers, vital for safe navigation and situational awareness.
4 papers · 1 benchmark
SWORD ('Scenes with occluded regions' dataset)
The new dataset contains around 1,500 train videos and 290 test videos, with 50 frames per video on average.
4 papers · 1 benchmark
SYMON (Synopses of Movie Narratives)
Contains 5,193 video summaries of popular movies and TV series.
4 papers · 0 benchmarks
SportsPose (SportsPose - A Dynamic 3D sports pose dataset)
Accurate 3D human pose estimation is essential for sports analytics, coaching, and injury prevention.
4 papers · 0 benchmarks
TinyVIRAT contains natural low-resolution activities.
4 papers · 0 benchmarks
Contains 140 videos with multiple human created summaries, which were acquired in a controlled experiment.
4 papers · 0 benchmarks
VOT2019 is a Visual Object Tracking benchmark for short-term tracking in RGB.
4 papers · 1 benchmark
The VideoNavQA dataset contains pairs of questions and videos generated in the House3D environment.
4 papers · 0 benchmarks
Youku-mPLUG is a large Chinese high-quality video-language dataset which is collected from Youku.com, a well-known Chinese video-sharing website, with strict criteria of safety, diversity, and quality.
4 papers · 0 benchmarks
The dataset is designed specifically to solve a range of computer vision problems (2D-3D tracking, posture) faced by biologists while designing behavior studies with animals.
3 papers · 0 benchmarks
AFEW-VA (AFEW-VA Database for Valence and Arousal Estimation In-The-Wild)
The AFEW-VA databaset is a collection of highly accurate per-frame annotations levels of valence and arousal, along with per-frame annotations of 68 facial landmarks for 600 challenging video clips.
3 papers · 0 benchmarks
ActivityNet Adverbs is a subset from the ActivityNet dataset with extracted verb-adverb annotations.
3 papers · 2 benchmarks
The Algonauts dataset provides human brain responses to a set of 1,102 3-s long video clips of everyday events.
3 papers · 0 benchmarks
BUAA-MIHR dataset is a remote photoplethysmography (rPPG) dataset.
3 papers · 0 benchmarks
CFC (Caltech Fish Counting Dataset)
Caltech Fish Counting Dataset (CFC) is a large-scale dataset for detecting, tracking, and counting fish in sonar videos.
3 papers · 0 benchmarks
The Composable activities dataset consists of 693 videos that contain activities in 16 classes performed by 14 actors.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.