Home › Datasets › task › Visual Tracking
Visual Tracking datasets
archive 2025-07-28
28 datasets carry the task tag "Visual Tracking" (the task itself: Visual Tracking), ordered by the archive's paper count. Page 1 of 1: 28 shown of 28. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Visual Tracking datasets 1–28 of 28
Kinetics (Kinetics Human Action Video Dataset)
The Kinetics dataset is a large-scale, high-quality dataset for human action recognition in videos.
1,341 papers · 18 benchmarks
DAVIS (Densely Annotated VIdeo Segmentation)
The Densely Annotation Video Segmentation dataset (DAVIS) is a high quality and high resolution densely annotated video segmentation dataset under two resolutions, 480p and 1080p.
734 papers · 10 benchmarks
Object Tracking Benchmark (OTB) is a visual tracking benchmark that is widely used to evaluate the performance of a visual tracking algorithm.
416 papers · 1 benchmark
LaSOT (Large-scale Single Object Tracking)
LaSOT is a high-quality benchmark for Large-scale Single Object Tracking.
275 papers · 3 benchmarks
TrackingNet is a large-scale tracking dataset consisting of videos in the wild.
210 papers · 2 benchmarks
OTB-2015, also referred as Visual Tracker Benchmark, is a visual tracking dataset.
182 papers · 1 benchmark
OTB2013 is the previous version of the current OTB2015 Visual Tracker Benchmark.
110 papers · 2 benchmarks
Kubric is a data generation pipeline for creating semi-realistic synthetic multi-object videos with rich annotations such as instance segmentation masks, depth maps, and optical flow.
77 papers · 1 benchmark
TNL2K (Tracking by natural language)
Tracking by Natural Language (TNL2K) is constructed for the evaluation of tracking by natural language specification.
62 papers · 2 benchmarks
TAP-Vid is a benchmark which contains both real-world videos with accurate human annotations of point tracks, and synthetic videos with perfect ground-truth point tracks.
46 papers · 1 benchmark
The Visual Object Tracking (VOT) dataset is a collection of video sequences used for evaluating and benchmarking visual object tracking algorithms.
36 papers · 0 benchmarks
OxUva is a dataset and benchmark for evaluating single-object tracking algorithms.
34 papers · 0 benchmarks
The Dialog State Tracking Challenges 2 & 3 (DSTC2&3) were research challenge focused on improving the state of the art in tracking the state of spoken dialog systems.
33 papers · 5 benchmarks
CDTB (Color-and-Depth Tracking)
Source: https://www.vicos.si/Projects/CDTB 4.2 State-of-the-art Comparison A TH CTB (color-and-depth visual object tracking) dataset is recorded by several passive and active RGB-D setups and contains indoor as well as outdoor sequences…
17 papers · 0 benchmarks
TLP (Track Long and Prosper)
A new long video dataset and benchmark for single object tracking.
14 papers · 0 benchmarks
RGB-Stacking is a benchmark for vision-based robotic manipulation.
13 papers · 3 benchmarks
VOT2014 (Visual Object Tracking Challenge 2014)
The dataset comprises 25 short sequences showing various objects in challenging backgrounds.
12 papers · 1 benchmark
YT-BB (YouTube-BoundingBoxes)
YouTube-BoundingBoxes (YT-BB) is a large-scale data set of video URLs with densely-sampled object bounding box annotations.
7 papers · 1 benchmark
MMPTRACK (Multi-camera Multiple People Tracking Dataset)
Multi-camera Multiple People Tracking (MMPTRACK) dataset has about 9.6 hours of videos, with over half a million frame-wise annotations.
6 papers · 1 benchmark
Description The consists of 92 groups of video clips with 113, 918 high resolution frames taken by two drones and 63 groups of video clips with 145, 875 high resolution frames taken by three drones.
4 papers · 0 benchmarks
SurgT is a dataset for benchmarking 2D Trackers in Minimally Invasive Surgery (MIS).
3 papers · 0 benchmarks
VBR (VBR: A Vision Benchmark in Rome)
This dataset presents a vision and perception research dataset collected in Rome, featuring RGB data, 3D point clouds, IMU, and GPS data.
3 papers · 0 benchmarks
Estimating camera motion in deformable scenes poses a complex and open research challenge.
2 papers · 1 benchmark
ARKitTrack is a new RGB-D tracking dataset for both static and dynamic scenes captured by consumer-grade LiDAR scanners equipped on Apple's iPhone and iPad.
1 paper · 0 benchmarks
LTFT (Long-Term Face Tracking)
Dataset originally conceived for multi-face tracking/detection for highly crowded scenarios.
1 paper · 0 benchmarks
MobiFace is the first dataset for single face tracking in mobile situations.
1 paper · 0 benchmarks
This dataset contains nine video sequences captured by a webcam for salient closed boundary tracking evaluation.
1 paper · 0 benchmarks
The dataset is composed of 100 video sequences densely annotated with 60K bounding boxes, 17 sequence attributes, 13 action verb attributes and 29 target object attributes.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.