Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 252 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 12049–12096 of 12,172
Description: The Traffic Sign Recognition Dataset is designed to support the development of deep learning models, particularly for object detection and classification.
0 papers · 0 benchmarks
An extended version of an experimental dataset, called Transient Biometrics Nails Dataset (TBND), was created.
0 papers · 0 benchmarks
This dataset is an extremely challenging set of over 3000+ original Transparent object images such as glasses and mirrors are captured and crowdsourced from over 500+ urban and rural areas, where each image is manually reviewed and…
0 papers · 0 benchmarks
The Trillion Word Corpus is a dataset created by Google, which contains one trillion words from public web pages.
0 papers · 0 benchmarks
UAV Trajectory (Simulated Dataset of UAV Food Delivery using Air Traffic Simulator)
The UAV Delievery dataset created to advance the research in drone delivery, contains trajectory details of UAV in different speed, altitude, and wind conditions.
0 papers · 0 benchmarks
UAV tracking dataset with common corruptions tailored for unmannered aerial video
0 papers · 0 benchmarks
UAS-based Multispectral othomosaics of vineyards from central Portugal - 2 distinct vineyards - Multispectral and HD orthomosaics
0 papers · 0 benchmarks
The UBIRIS.v2 iris dataset contains 11,102 iris images from 261 subjects with 10 images each subject.
0 papers · 0 benchmarks
The Tagalog Universal Dependencies NewsCrawl dataset consists of annotated text extracted from the Leipzig Tagalog Corpus.
0 papers · 0 benchmarks
UFO Cherry Tree Point Clouds consists of a collection of 82 scanned Upright Fruiting Offshoot (UFO) cherry tree point clouds.
0 papers · 0 benchmarks
The University of Massachusetts Amherst citation field extraction dataset contains labels and segments for extracted citations from articles found on arXiv.
0 papers · 0 benchmarks
UNIPD-BPE (University of Padova Body Pose Estimation)
The University of Padova Body Pose Estimation dataset (UNIPD-BPE) is an extensive dataset for multi-sensor body pose estimation containing both single-person and multi-person sequences with up to 4 interacting people A network with 5…
0 papers · 0 benchmarks
UQ NIDS (ML-based NIDS NetFlow Datasets)
The datasets on this page are designed for machine learning-based Network Intrusion Detection Systems (NIDS) and are organised into the following high-level collections: NetFlow V3 Datasets: This collection consists of four datasets in…
0 papers · 0 benchmarks
USYD CAMPUS is a driving dataset collected by Zhou et al at the University of Sydney (USyd) campus and surroundings.
0 papers · 0 benchmarks
UT Zappos50K (UT-Zap50K) is a large shoe dataset consisting of 50,025 catalog images collected from Zappos.com.
0 papers · 0 benchmarks
Dataset for range data gathered using the Ultra-wide Band (UWB) MDEK 1001 Dev.
0 papers · 0 benchmarks
The Unsplash Dataset is created by over 200,000 contributing photographers and billions of searches across thousands of applications, uses, and contexts.
0 papers · 0 benchmarks
Urdu Sentiment Corpus (Urdu Sentiment Corpus (v1.0): Linguistic Exploration and Visualization of Labeled Datasetfor Urdu Sentiment Analysis)
Consists of Urdu tweets for the sentiment analysis and polarity detection.
0 papers · 0 benchmarks
The VICAVR database is a set of retinal images used for the computation of the A/V Ratio.
0 papers · 0 benchmarks
Human activity recognition and clinical biomechanics are challenging problems in physical telerehabilitation medicine.
0 papers · 0 benchmarks
The dataset contains more than 35000 images and 600 videos captured using 35 different portable devices of 11 major brands.
0 papers · 0 benchmarks
VOT2013 (Visual Object Tracking Challenge 2013)
The dataset comprises 16 short sequences showing various objects in challenging backgrounds.
0 papers · 0 benchmarks
To facilitate the research of vehicle re-identification (Re-Id), a large-scale benchmark dateset is built for vehicle Re-Id in the real-world urban surveillance scenario, named “VeRi”.
0 papers · 0 benchmarks
This dataset is collected by DataCluster Labs.
0 papers · 0 benchmarks
Vehicle-1M involves vehicle images captured across day and night, from head or rear, by multiple surveillance cameras installed in cities.
0 papers · 0 benchmarks
ViDoSeek, a benchmark specifically designed for visually rich document retrieval-reason-answer, fully suited for evaluation of RAG within large document corpus.
0 papers · 0 benchmarks
This dataset contains over 12K+ questions and symptoms related to various common diseases in Vietnamese.
0 papers · 0 benchmarks
we construct the ViMirr dataset, which has 19,255 frames from 276 videos
0 papers · 0 benchmarks
Video Dataset (Storytelling Video Dataset (Russian, Emotion, Gesture, Speech))
The Storytelling Video Dataset is a high-quality, human-reviewed multimodal dataset featuring over 700 full-body video recordings of native Russian speakers.
0 papers · 0 benchmarks
Crowd Violence \ Non-violence Database and benchmark: A database of real-world, video footage of crowd violence, along with standard benchmark protocols designed to test both violent/non-violent classification and violence outbreak…
0 papers · 0 benchmarks
This dataset is an extremely challenging set of over 2000+ original Visiting card/ID card images captured and crowdsourced from over 300+ urban and rural areas, where each image is manually reviewed and verified by computer vision…
0 papers · 0 benchmarks
Visual Fields (UWHVF: A real-world, open source dataset of Humphrey Visual Fields (HVF) from the University of Washington)
28,943 Humphrey Visual Field (HVF) tests from 3,871 patients and 7,428 eyes.
0 papers · 0 benchmarks
VocSim (Vocal Similarity Benchmark)
VocSim (Vocal Similarity Benchmark) is a benchmark designed to evaluate the ability of neural audio embeddings to capture acoustic and perceptual similarity in a zero-shot setting, without task-specific fine-tuning.
0 papers · 0 benchmarks
This dataset is part of my bachelor thesis project.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.