Home › Datasets › modality › Tracking
Tracking datasets
archive 2025-07-28
77 datasets carry the modality tag "Tracking", ordered by the archive's paper count. Page 1 of 2: 48 shown of 77. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Tracking datasets 1–48 of 77
The nuScenes dataset is a large-scale autonomous driving dataset.
2,139 papers · 21 benchmarks
Object Tracking Benchmark (OTB) is a visual tracking benchmark that is widely used to evaluate the performance of a visual tracking algorithm.
416 papers · 1 benchmark
TrackingNet is a large-scale tracking dataset consisting of videos in the wild.
210 papers · 2 benchmarks
highD Dataset (The Highway Drone Dataset Naturalistic Trajectories of 110 500 Vehicles Recorded at German Highways)
The highD dataset is a new dataset of naturalistic vehicle trajectories recorded on German highways.
115 papers · 0 benchmarks
VOT2016 is a video dataset for visual object tracking.
113 papers · 1 benchmark
The PoseTrack dataset is a large-scale benchmark for multi-person pose estimation and tracking in videos.
103 papers · 5 benchmarks
MOT15 (Multiple Object Tracking 15)
MOT2015 is a dataset for multiple object tracking.
67 papers · 5 benchmarks
VOT2017 (Visual Object Tracking Challenge)
VOT2017 is a Visual Object Tracking dataset for different tasks that contains 60 short sequences annotated with 6 different attributes.
56 papers · 2 benchmarks
The inD dataset is a new dataset of naturalistic vehicle trajectories recorded at German intersections.
49 papers · 0 benchmarks
The RGBT234 dataset is a comprehensive video dataset specifically designed for RGB-T (Red-Green-Blue and Thermal) tracking purposes.
38 papers · 1 benchmark
KITTI MOTS (KITTI Multi-Object Tracking and Segmentation (MOTS) Evaluation)
The Multi-Object and Segmentation (MOTS) benchmark [2] consists of 21 training sequences and 29 test sequences.
28 papers · 1 benchmark
The rounD dataset introduces a fresh compilation of natural road user trajectory data from German roundabouts, gathered using drone technology to navigate past usual challenges such as occlusions inherent in traditional traffic data…
19 papers · 0 benchmarks
Expi (Extreme Pose Interaction)
Extreme Pose Interaction (ExPI) Dataset is a new person interaction dataset of Lindy Hop dancing actions.
16 papers · 3 benchmarks
VOT2014 (Visual Object Tracking Challenge 2014)
The dataset comprises 25 short sequences showing various objects in challenging backgrounds.
12 papers · 1 benchmark
REFLACX (Reports and eye-tracking data for localization of abnormalities in chest x-rays)
The REFLACX dataset contains eye-tracking data for 3,032 readings of chest x-rays by five radiologists.
10 papers · 0 benchmarks
Atari-HEAD is a dataset of human actions and eye movements recorded while playing Atari videos games.
9 papers · 0 benchmarks
We provide manual annotations of 14 semantic keypoints for 100,000 car instances (sedan, suv, bus, and truck) from 53,000 images captured from 18 moving cameras at Multiple intersections in Pittsburgh, PA.
9 papers · 2 benchmarks
PathTrack is a dataset for person tracking which contains more than 15,000 person trajectories in 720 sequences.
9 papers · 0 benchmarks
Most existing MOT datasets are captured using pinhole cameras, which are characterized by a narrow-FoV and linear sensor motion.
9 papers · 1 benchmark
VOT2020 is a Visual Object Tracking benchmark for short-term tracking in RGB.
9 papers · 1 benchmark
WALT (Watch and Learn TimeLapse Images)
We introduce a new dataset, Watch and Learn Time-lapse (WALT), consisting of multiple (4K and 1080p) cameras capturing urban environments over a year.
7 papers · 1 benchmark
300-VW (300 Videos in the Wild)
300 Videos in the Wild (300-VW) is a dataset for evaluating facial landmark tracking algorithms in the wild.
6 papers · 2 benchmarks
IndustReal (IndustReal Dataset of Egocentric Videos for Procedure Understanding)
IndustReal is an ego-centric, multi-modal dataset where 27 participants are challenged to perform assembly and maintenance procedures on a construction-toy car.
6 papers · 3 benchmarks
MMPTRACK (Multi-camera Multiple People Tracking Dataset)
Multi-camera Multiple People Tracking (MMPTRACK) dataset has about 9.6 hours of videos, with over half a million frame-wise annotations.
6 papers · 1 benchmark
A new dataset with significant occlusions related to object manipulation.
6 papers · 0 benchmarks
The exiD dataset introduces a groundbreaking collection of naturalistic road user trajectories at highway entries and exits in Germany, meticulously captured with drones to navigate past the limitations of conventional traffic data…
6 papers · 0 benchmarks
aiMotive dataset is a multimodal dataset for robust autonomous driving with long-range perception.
5 papers · 1 benchmark
VOT2019 is a Visual Object Tracking benchmark for short-term tracking in RGB.
4 papers · 1 benchmark
The eSports Sensors dataset contains sensor data collected from 10 players in 22 matches in League of Legends.
4 papers · 2 benchmarks
The dataset is designed specifically to solve a range of computer vision problems (2D-3D tracking, posture) faced by biologists while designing behavior studies with animals.
3 papers · 0 benchmarks
DivEMT (Post-Editing Effort Across Typologically-diverse Languages)
DivEMT, the first publicly available post-editing study of Neural Machine Translation (NMT) over a typologically diverse set of target languages.
3 papers · 0 benchmarks
GroOT (Grounded Multiple Object Tracking)
One of the recent trends in vision problems is to use natural language captions to describe the objects of interest.
3 papers · 0 benchmarks
The SoccerNet Game State Reconstruction task is a novel high level computer vision task that is specific to sports analytics.
3 papers · 0 benchmarks
The UAVA,UAV-Assistant, dataset is specifically designed for fostering applications which consider UAVs and humans as cooperative agents.
3 papers · 0 benchmarks
VBR (VBR: A Vision Benchmark in Rome)
This dataset presents a vision and perception research dataset collected in Rome, featuring RGB data, 3D point clouds, IMU, and GPS data.
3 papers · 0 benchmarks
CVB (Video Dataset of Cattle Visual Behaviors)
Existing image/video datasets for cattle behavior recognition are mostly small, lack well-defined labels, or are collected in unrealistic controlled environments.
2 papers · 0 benchmarks
The NBA SportVU dataset contains player and ball trajectories for 631 games from the 2015-2016 NBA season.
2 papers · 1 benchmark
This dataset comprehends the 3D building information model (in IFC and Revit formats), manually elaborated based on the terrestrial laser scanner of the sequence 2 of ConSLAM, and the refined ground truth (GT) poses (in TUM format) of…
2 papers · 0 benchmarks
SNDZoo (The Softwarised Network Data Zoo)
The softwarised network data zoo (SNDZoo) is an open collection of software networking data sets aiming to streamline and ease machine learning research in the software networking domain.
2 papers · 0 benchmarks
Data used for the paper SparsePoser: Real-time Full-body Motion Reconstruction from Sparse Data It contains over 1GB of high-quality motion capture data recorded with an Xsens Awinda system while using a variety of VR applications in Meta…
2 papers · 0 benchmarks
This is a comprehensive dataset of human arm motion during Activities of Daily Living (ADL).
1 paper · 0 benchmarks
AViMoS (Audio-Visual Mouse Saliency)
A novel audio-visual mouse saliency (AViMoS) dataset with the following key-features: Diverse content: movie, sports, live, vertical videos, etc.; Large scale: 1500 videos with mean 19s duration; High resolution: all streams are FullHD;…
1 paper · 0 benchmarks
Temporal Dataset for Indoor and In-Vehicle Thermal Comfort Estimation Abstract Thermal comfort estimation is essential for enhancing user experience in static indoor environments and dynamic in-vehicle scenarios.
1 paper · 0 benchmarks
We present a new simulated dataset for pedestrian action anticipation collected using the CARLA simulator.
1 paper · 0 benchmarks
This dataset contains Axivity AX3 wrist-worn activity tracker data that were collected from 151 participants in 2014-2016 around the Oxfordshire area.
1 paper · 0 benchmarks
ConSLAM (Construction Dataset for SLAM)
ConSLAM is a real-world dataset collected periodically on a construction site to measure the accuracy of mobile scanners' SLAM algorithms.
1 paper · 0 benchmarks
Data for "Image-based Backbone Reconstruction for Non-Slender Soft Robots" This dataset provides the data for the forthcoming paper "Image-based Backbone Reconstruction for Non-Slender Soft Robots".
1 paper · 0 benchmarks
Collected data from two distinct experiments in immersive, interactive VR where participants performed dynamic tasks as their eye, head, and hand movements were recorded.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.