Home › Datasets › task › Video Inpainting
Video Inpainting datasets
archive 2025-07-28
15 datasets carry the task tag "Video Inpainting" (the task itself: Video Inpainting), ordered by the archive's paper count. Page 1 of 1: 15 shown of 15. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Video Inpainting datasets 1–15 of 15
DAVIS (Densely Annotated VIdeo Segmentation)
The Densely Annotation Video Segmentation dataset (DAVIS) is a high quality and high resolution densely annotated video segmentation dataset under two resolutions, 480p and 1080p.
734 papers · 10 benchmarks
Youtube-VOS is a Video Object Segmentation dataset that contains 4,453 videos - 3,471 for training, 474 for validation, and 508 for testing.
203 papers · 10 benchmarks
How2Sign (A Large-scale Multimodal Dataset for Continuous American Sign Language)
The How2Sign is a multimodal and multiview continuous American Sign Language (ASL) dataset consisting of a parallel corpus of more than 80 hours of sign language videos and a set of corresponding modalities including speech, English…
44 papers · 3 benchmarks
EgoHOS (Fine-Grained Egocentric Hand-Object Segmentation Dataset)
EgoHOS is a labeled dataset consisting of 11243 egocentric images with per-pixel segmentation labels of hands and objects being interacted with during a diverse array of daily activities.
9 papers · 0 benchmarks
Are current 3D object tracking methods truely robust enough for low-fidelity depth sensors like the iPhone LiDAR?
8 papers · 2 benchmarks
FVI (Free-form Video Inpainting)
The Free-Form Video Inpainting dataset is a dataset used for training and evaluation video inpainting models.
8 papers · 0 benchmarks
KITTI360-EX is a dataset for outer- and inner FoV expansion.
7 papers · 1 benchmark
DEVIL (Diagnostic Evaluation of Video Inpainting on Landscapes)
Diagnostic Evaluation of Video Inpainting on Landscapes (DEVIL) benchmark is composed of a curated video/occlusion mask dataset and a comprehensive evaluation scheme
3 papers · 0 benchmarks
QST contains 1,167 video clips that are cut out from 216 time-lapse 4K videos collected from YouTube, which can be used for a variety of tasks, such as (high-resolution) video generation, (high-resolution) video prediction,…
3 papers · 0 benchmarks
The Inpainting dataset consists of synchronized Labeled image and LiDAR scanned point clouds.
1 paper · 1 benchmark
The benchmark for VPData, the largest video inpainting dataset, which comprises over 390K clips (> 866.7 hours) and features precise masks and detailed video captions.
1 paper · 0 benchmarks
The largest video inpainting dataset comprises over 390K clips (> 866.7 hours), featuring precise masks and detailed video captions.
1 paper · 0 benchmarks
We provide video sequences with annotated object masks for video inpainting.
1 paper · 0 benchmarks
Dataset for the DREAMING - Diminished Reality for Emerging Applications in Medicine through Inpainting Challenge!
0 papers · 0 benchmarks
WRV (Wire-removal Dataset)
G2LP Wire-removal Dataset in G2LP-Net: Global to Local Progressive Video Inpainting Network Wire-removal Dataset The WRV dataset has been specifically curated for the challenges of video inpainting in irregularly slender regions.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.