Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 223 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 10657–10704 of 12,172
Definitions of jargon/terms in computer science, mathematics, and physics
1 paper · 0 benchmarks
This package contains an anonymized packets of 802.11 probe requests captured throughout March of 2023 at Universitat Jaume I.
1 paper · 0 benchmarks
The UJIIndoorLoc is a Multi-Building Multi-Floor indoor localization database to test Indoor Positioning System that rely on WLAN/WiFi fingerprint.
1 paper · 0 benchmarks
Education is increasingly data-driven, and the ability to analyse and adapt educational materials quickly and effectively is important for keeping materials contemporary and interesting.
1 paper · 1 benchmark
Bangladesh's legal system struggles with major challenges like delays, complexity, high costs, and millions of unresolved cases, which deter many from pursuing legal action due to lack of knowledge or financial constraints.
1 paper · 0 benchmarks
ULI-RI (Unreal Labeled Images for Person Re-ID)
The ULI-RI dataset is generated using the Unreal Engine 4 to simulate various outdoor environments with 115 high-quality 3D human models.
1 paper · 0 benchmarks
ULS labeled data (UVA laser scanning labelled las data over tropical moist forest classified as leaf or wood points)
UAV Laser Scanning data collected over neotropical forest (Paracou French Guiana).
1 paper · 1 benchmark
UMA-VI Dataset (The UMA-VI dataset: Visual--inertial odometry in low-textured and dynamic illumination environments)
The dataset contains 32 sequences for the evaluation of VI motion estimation methods, totalling ∼80 min of data.
1 paper · 0 benchmarks
UMAD (Urban Minimum Altitude Dataset)
The UMAD is a virtual-scene dataset made by AirSim, which is a simulator built on Unreal Engine.
1 paper · 0 benchmarks
UMC005 English-Urdu is a parallel corpus of texts in English and Urdu language with sentence alignments.
1 paper · 0 benchmarks
One-Shot Affordance Part Segmentation variant of the UMD dataset.
1 paper · 0 benchmarks
Repository for UML-English data This repository contains the data used for "Extraction of UML Class Diagrams from Natural Language Specification" (Yang et al.
1 paper · 1 benchmark
UNER v1 adds an NER annotation layer to 18 datasets (primarily treebanks from UD) and covers 12 geneologically and ty- pologically diverse languages: Cebuano, Danish, German, English, Croatian, Portuguese, Russian, Slovak, Serbian,…
1 paper · 31 benchmarks
A comprehensive dataset, merging all the aforementioned datasets.
1 paper · 0 benchmarks
CICFlowMeter format of the datasets are made up of 83 features.
1 paper · 0 benchmarks
A comprehensive dataset, merging all the aforementioned datasets.
1 paper · 0 benchmarks
In this dataset UR5 robot used 6 tools: metal-scissor, metal-whisk, plastic-knife, plastic-spoon, wooden-chopstick, and wooden-fork to perform 6 behaviors: look, stirring-slow, stirring-fast, stirring-twist, whisk, and poke.
1 paper · 0 benchmarks
This is a countrywide traffic accident dataset, which covers 49 states of the United States.
1 paper · 0 benchmarks
USB (Universal-Scale Object Detection Benchmark)
The Universal-Scale object detection Benchmark (USB) is a benchmark for object detection that has variations in object scales and image domains by incorporating COCO with the recently proposed Waymo Open Dataset and Manga109-s dataset.
1 paper · 1 benchmark
USC (Uzbek Speech Corpus)
The Uzbek speech corpus (USC) comprises 958 different speakers with a total of 105 hours of transcribed audio recordings.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
USC-GRAD-STDdb comprises 115 video segments containing more than 25,000 annotated frames of HD 720p resolution (≈1280x720) with small objects of interest from 16 (≈4x4) to 256 (≈16x16) as pixel area.
1 paper · 1 benchmark
USCOCO (Unexpected Situations of Common Objects in Context)
A test set of grammatically correct sentences and layouts (visual “imagined” situations), called Unexpected Situations of Common Objects in Context (USCOCO) describing compositions of entities and relations that are unlikely to be found in…
1 paper · 0 benchmarks
Landsat 8 Collection 1 Tier 1 and Real-Time data DN values, representing scaled, calibrated at-sensor radiance.
1 paper · 0 benchmarks
USM-SED is a dataset for polyphonic sound event detection in urban sound monitoring use-cases.
1 paper · 0 benchmarks
The USNA long-term scintillation study is a continuing effort to characterize and measure optical turbulence in the near-maritime boundary layer.
1 paper · 1 benchmark
The USNA long-term scintillation study is a continuing effort to characterize and measure optical turbulence in the near-maritime boundary layer.
1 paper · 2 benchmarks
The USPTO Backgrounds dataset provides valuable information related to patents and trademarks.
1 paper · 1 benchmark
We introduce USPTO-30K, a large-scale benchmark dataset of annotated molecule images, which overcomes these limitations.
1 paper · 0 benchmarks
UT-GLOBUS (GLObal Building heights for Urban Studies)
We introduce GLObal Building heights for Urban Studies (UT-GLOBUS), a dataset providing building heights and urban canopy parameters (UCPs) for major cities worldwide.
1 paper · 0 benchmarks
UTCD (Universal Text Classification Dataset)
UTCD is a compilation of 18 classification datasets spanning 3 categories of Sentiment, Intent/Dialogue, and Topic classification.
1 paper · 0 benchmarks
The semantic segmentation of clothes is a challenging task due to the wide variety of clothing styles, layers and shapes.
1 paper · 1 benchmark
This 2d indoor dataset collection consists of 9 individual datasets.
1 paper · 0 benchmarks
The UTRSet-Real dataset is a comprehensive, manually annotated dataset specifically curated for Printed Urdu OCR research.
1 paper · 0 benchmarks
The UTRSet-Synth dataset is introduced as a complementary training resource to the UTRSet-Real Dataset, specifically designed to enhance the effectiveness of Urdu OCR models.
1 paper · 0 benchmarks
UTSig (UTSig: A Persian offline signature dataset)
UTSig (University of Tehran Persian Signature) dataset is freely available at MLCM lab website: http://mlcm.ut.ac.ir/Datasets.html Persian offline signature dataset, UTSig.
1 paper · 0 benchmarks
UV6K (Urban Vehicle Segmentation Dataset)
UV6K is a high-resolution remote sensing urban vehicle segmentation dataset.
1 paper · 1 benchmark
UW IOM (University of Washington Indoor Object Manipulation)
Comprises twenty individuals picking up and placing objects of varying weights to and from cabinet and table locations at various heights.
1 paper · 0 benchmarks
The Ubuntu Chat Corpus (UCC) is composed of archived chat logs from Ubuntu's Internet Relay Chat technical support channels.
1 paper · 0 benchmarks
a character sheet dataset containing over 700,000 hand-drawn and synthesized images of diverse poses
1 paper · 0 benchmarks
UltraLAMBDAis a large-scale dataset of ads sourced from brand videos on platforms such as YouTube and Facebook Ads, as well as from CommonCrawl.
1 paper · 0 benchmarks
Street-View images captured at different timestamps often undergo geometric transformations.
1 paper · 1 benchmark
This dataset extends the Semantic Segmentation of Underwater Imagery: Dataset and Benchmark, adding an uncertainty evaluation component.
1 paper · 0 benchmarks
AI-based digital twins are at the leading edge of theIndustry 4.0 revolution, which are technologically empowered bythe Internet of Things and real-time data analysis.
1 paper · 0 benchmarks
Uncorrelated Corrupted Dataset is an evaluation set that consists of realistic visible-infrared (V-I) corruptions allowing for models' corruption robustness evaluation.
1 paper · 0 benchmarks
This data contains the election polls for the 2004, 2008, 2012, and 2016 US presidential election by state including data on undecided voter proportions.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
We construct a large-scale Heterogeneous Graph benchmark dataset named UniKG from Wikidata.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.