Home › Datasets › task › Video Object Detection

Video Object Detection datasets

archive 2025-07-28

13 datasets carry the task tag "Video Object Detection" (the task itself: Video Object Detection), ordered by the archive's paper count. Page 1 of 1: 13 shown of 13. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Video Object Detection datasets 1–13 of 13

The Waymo Open Dataset is comprised of high resolution sensor data collected by autonomous vehicles operated by the Waymo Driver in a wide variety of conditions.
481 papers · 16 benchmarks
The EPIC-KITCHENS-55 dataset comprises a set of 432 egocentric videos recorded by 32 participants in their kitchens at 60fps with a head mounted camera.
42 papers · 3 benchmarks
ImageNet VID is a large-scale public dataset for video object detection and contains more than 1M frames for training and more than 100k frames for validation.
26 papers · 1 benchmark
GEN1 Detection (Prophesee GEN1 Automotive Detection Dataset)
Prophesee’s GEN1 Automotive Detection Dataset is the largest Event-Based Dataset to date.
13 papers · 1 benchmark
Are current 3D object tracking methods truely robust enough for low-fidelity depth sensors like the iPhone LiDAR?
8 papers · 2 benchmarks
YT-BB (YouTube-BoundingBoxes)
YouTube-BoundingBoxes (YT-BB) is a large-scale data set of video URLs with densely-sampled object bounding box annotations.
7 papers · 1 benchmark
OAK (Objects Around Krishna)
OAK is a dataset for online continual object detection benchmark with an egocentric video dataset.
5 papers · 0 benchmarks
Specially designed to evaluate active learning for video object detection in road scenes.
5 papers · 0 benchmarks
Description: 5,011 Images – Human Frontal face Data (Male).
2 papers · 0 benchmarks
Underwater Trash Detection Dataset Overview The Underwater Trash Detection Dataset is a custom-annotated dataset designed to address the challenges of underwater trash detection caused by varying environmental features.
2 papers · 0 benchmarks
THGP (Temporal Hands Guns and Phones Dataset)
Temporal Hands Guns and Phones (THGP) dataset, is a collection of 5960 video frames (5000 for training and 960 for testing).
1 paper · 0 benchmarks
USC-GRAD-STDdb (Small Target Detection database)
USC-GRAD-STDdb comprises 115 video segments containing more than 25,000 annotated frames of HD 720p resolution (≈1280x720) with small objects of interest from 16 (≈4x4) to 256 (≈16x16) as pixel area.
1 paper · 1 benchmark
VISEM-Tracking is a dataset consisting of 20 video recordings of 30s of spermatozoa with manually annotated bounding-box coordinates and a set of sperm characteristics analyzed by experts in the domain.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.