Home › Datasets › task › Video Segmentation
Video Segmentation datasets
archive 2025-07-28
12 datasets carry the task tag "Video Segmentation" (the task itself: Video Segmentation), ordered by the archive's paper count. Page 1 of 1: 12 shown of 12. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Video Segmentation datasets 1–12 of 12
Datasets drive vision progress, yet existing driving datasets are impoverished in terms of visual content and supported tasks to study multitask learning for autonomous driving.
469 papers · 16 benchmarks
SegTrack v2 is a video segmentation dataset with full pixel-level annotations on multiple objects at each frame within each video.
107 papers · 5 benchmarks
Dynamic Replica is a synthetic dataset of stereo videos featuring humans and animals in virtual environments.
10 papers · 0 benchmarks
EgoProceL is a large-scale dataset for procedure learning.
9 papers · 0 benchmarks
TikTok Dataset (Learning High Fidelity Depths of Dressed Humans by Watching Social Media Dance Videos)
We learn high fidelity human depths by leveraging a collection of social media dance videos scraped from the TikTok mobile social networking application.
8 papers · 0 benchmarks
Operating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assistance, situational awareness, and patient safety.
3 papers · 3 benchmarks
MOMA-LRG (Multi-Object Multi-Actor activity parsing with Language-Refined Graphs)
A dataset dedicated to multi-object, multi-actor activity parsing.
3 papers · 1 benchmark
PETRAW (PEg TRAnsfer Workflow recognition by different modalities)
PETRAW data set was composed of 150 sequences of peg transfer training sessions.
3 papers · 6 benchmarks
A large-scale video portrait dataset that contains 291 videos from 23 conference scenes with 14K frames.
3 papers · 0 benchmarks
This is a video and image segmentation dataset for human head and shoulders, relevant for creating elegant media for videoconferencing and virtual reality applications.
1 paper · 0 benchmarks
LSDBench (Long-video Sampling Dilemma Benchmark)
A benchmark that focuses on the sampling dilemma in long-video tasks.
1 paper · 0 benchmarks
Infinity AI's Spills Basic Dataset is a synthetic, open-source dataset for safety applications.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.