Home › Datasets › task › Object Tracking
Object Tracking datasets
archive 2025-07-28
69 datasets carry the task tag "Object Tracking" (the task itself: Object Tracking), ordered by the archive's paper count. Page 1 of 2: 48 shown of 69. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Object Tracking datasets 1–48 of 69
KITTI (Karlsruhe Institute of Technology and Toyota Technological Institute) is one of the most popular datasets for use in mobile robotics and autonomous driving.
3,661 papers · 137 benchmarks
LaSOT (Large-scale Single Object Tracking)
LaSOT is a high-quality benchmark for Large-scale Single Object Tracking.
275 papers · 3 benchmarks
GOT-10k (Generic Object Tracking Benchmark)
The GOT-10k dataset contains more than 10,000 video segments of real-world moving objects and over 1.5 million manually labelled bounding boxes.
239 papers · 2 benchmarks
The MOTChallenge datasets are designed for the task of multiple object tracking.
192 papers · 0 benchmarks
Virtual KITTI is a photo-realistic synthetic video dataset designed to learn and evaluate computer vision models for several video understanding tasks: object detection and multi-object tracking, scene-level and instance-level semantic…
133 papers · 0 benchmarks
VOT2018 is a dataset for visual object tracking.
129 papers · 1 benchmark
UAVDT (Unmanned Aerial Vehicle Benchmark Object Detection and Tracking)
UAVDT is a large scale challenging UAV Detection and Tracking benchmark (i.e., about 80, 000 representative frames from 10 hours raw videos) for 3 important fundamental tasks, i.e., object DETection (DET), Single Object Tracking (SOT) and…
96 papers · 2 benchmarks
VOT2017 (Visual Object Tracking Challenge)
VOT2017 is a Visual Object Tracking dataset for different tasks that contains 60 short sequences annotated with 6 different attributes.
56 papers · 2 benchmarks
Consists of 100 challenging video sequences captured from real-world traffic scenes (over 140,000 frames with rich annotations, including occlusion, weather, vehicle category, truncation, and vehicle bounding boxes) for object detection,…
53 papers · 2 benchmarks
Virtual KITTI 2 is an updated version of the well-known Virtual KITTI dataset which consists of 5 sequence clones from the KITTI tracking benchmark.
53 papers · 2 benchmarks
The Event-Camera Dataset is a collection of datasets with an event-based camera for high-speed robotics.
51 papers · 2 benchmarks
TAO (Tracking Any Object Dataset)
TAO is a federated dataset for Tracking Any Object, containing 2,907 high resolution videos, captured in diverse environments, which are half a minute long on average.
49 papers · 1 benchmark
The Objectron dataset is a collection of short, object-centric video clips, which are accompanied by AR session metadata that includes camera poses, sparse point-clouds and characterization of the planar surfaces in the surrounding…
48 papers · 0 benchmarks
The Visual Object Tracking (VOT) dataset is a collection of video sequences used for evaluating and benchmarking visual object tracking algorithms.
36 papers · 0 benchmarks
OxUva is a dataset and benchmark for evaluating single-object tracking algorithms.
34 papers · 0 benchmarks
KITTI MOTS (KITTI Multi-Object Tracking and Segmentation (MOTS) Evaluation)
The Multi-Object and Segmentation (MOTS) benchmark [2] consists of 21 training sequences and 29 test sequences.
28 papers · 1 benchmark
A new video dataset for aerial view concurrent human action detection.
24 papers · 1 benchmark
VisEvent (Visible-Event benchmark) is a dataset constructed for the evaluation of tracking by combing visible and event cameras.
23 papers · 1 benchmark
A new large-scale dataset for understanding human motions, poses, and actions in a variety of realistic events, especially crowd & complex events.
19 papers · 1 benchmark
SeaDronesSee (SeaDronesSee: A Maritime Benchmark for Detecting Humans in Open Water)
SeaDronesSee is a large-scale data set aimed at helping develop systems for Search and Rescue (SAR) using Unmanned Aerial Vehicles (UAVs) in maritime scenarios.
18 papers · 3 benchmarks
CDTB (Color-and-Depth Tracking)
Source: https://www.vicos.si/Projects/CDTB 4.2 State-of-the-art Comparison A TH CTB (color-and-depth visual object tracking) dataset is recorded by several passive and active RGB-D setups and contains indoor as well as outdoor sequences…
17 papers · 0 benchmarks
This dataset includes 4,500 fully annotated images (over 30,000 license plate characters) from 150 vehicles in real-world scenarios where both the vehicle and the camera (inside another vehicle) are moving.
15 papers · 1 benchmark
TLP (Track Long and Prosper)
A new long video dataset and benchmark for single object tracking.
14 papers · 0 benchmarks
Large-scale single-object tracking dataset, containing 108 sequences with a total length of 1.5 hours.
13 papers · 1 benchmark
VOT2014 (Visual Object Tracking Challenge 2014)
The dataset comprises 25 short sequences showing various objects in challenging backgrounds.
12 papers · 1 benchmark
The AI City Challenge, hosted at CVPR 2024, focuses on harnessing AI to enhance operational efficiency in physical settings such as retail and warehouse environments, and Intelligent Traffic Systems (ITS).
10 papers · 1 benchmark
Perception Test is a benchmark designed to evaluate the perception and reasoning skills of multimodal models.
10 papers · 3 benchmarks
In this work, we propose a general dataset for Color-Event camera based Single Object Tracking, termed COESOT.
9 papers · 1 benchmark
PTB-TIR is a Thermal InfraRed (TIR) pedestrian tracking benchmark, which provides 60 TIR sequences with mannuly annoations.
9 papers · 0 benchmarks
PathTrack is a dataset for person tracking which contains more than 15,000 person trajectories in 720 sequences.
9 papers · 0 benchmarks
Most existing MOT datasets are captured using pinhole cameras, which are characterized by a narrow-FoV and linear sensor motion.
9 papers · 1 benchmark
DOLPHINS (Dataset for Collaborative Perception enabled Harmonious and Interconnected Self-driving)
Vehicle-to-Everything (V2X) network has enabled collaborative perception in autonomous driving, which is a promising solution to the fundamental defect of stand-alone intelligence including blind zones and long-range perception.
7 papers · 0 benchmarks
TREK-150 is a benchmark dataset for object tracking in First Person Vision (FPV) videos composed of 150 densely annotated video sequences.
7 papers · 0 benchmarks
MMPTRACK (Multi-camera Multiple People Tracking Dataset)
Multi-camera Multiple People Tracking (MMPTRACK) dataset has about 9.6 hours of videos, with over half a million frame-wise annotations.
6 papers · 1 benchmark
VideoCube is a high-quality and large-scale benchmark to create a challenging real-world experimental environment for Global Instance Tracking (GIT).
6 papers · 1 benchmark
The evaluation of object detection models is usually performed by optimizing a single metric, e.g.
5 papers · 1 benchmark
aiMotive dataset is a multimodal dataset for robust autonomous driving with long-range perception.
5 papers · 1 benchmark
HOMER (Household Object Movements from Everyday Routines)
The Household Object Movements from Everyday Routines (HOMER) dataset is composed of routine behaviors for five households, spanning 50 days for the train split and 10 days for test split.
4 papers · 0 benchmarks
Description The consists of 92 groups of video clips with 113, 918 high resolution frames taken by two drones and 63 groups of video clips with 145, 875 high resolution frames taken by three drones.
4 papers · 0 benchmarks
VOT2019 is a Visual Object Tracking benchmark for short-term tracking in RGB.
4 papers · 1 benchmark
WISDOM (Warehouse Instance Segmentation Dataset for Object Manipulation)
Synthetic training dataset of 50,000 depth images and 320,000 object masks using simulated heaps of 3D CAD models.
4 papers · 1 benchmark
The dataset is designed specifically to solve a range of computer vision problems (2D-3D tracking, posture) faced by biologists while designing behavior studies with animals.
3 papers · 0 benchmarks
CFC (Caltech Fish Counting Dataset)
Caltech Fish Counting Dataset (CFC) is a large-scale dataset for detecting, tracking, and counting fish in sonar videos.
3 papers · 0 benchmarks
DIVOTrack is a cross-view multi-object tracking dataset for DIVerse Open scenes with dense tracking pedestrians in realistic and non-experimental environments.
3 papers · 0 benchmarks
PRED18 (PRED18: Predator/Prey DAVIS Dataset)
Twenty DAVIS recordings with a total duration of about 1.25 hour were obtained by driving the two robots in the robot arena of the University of Ulster in Londonderry.
3 papers · 0 benchmarks
UAV-GESTURE is a dataset for UAV control and gesture recognition.
3 papers · 0 benchmarks
The AU-AIR is a multi-modal aerial dataset captured by a UAV.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.