Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 147 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 7009–7056 of 12,172
Tilde MODEL Corpus is a multilingual corpora for European languages – particularly focused on the smaller languages.
2 papers · 0 benchmarks
Tiny ImageNet-R is a subset of the ImageNet-R dataset by Hendrycks et al.
2 papers · 0 benchmarks
Titanic (Titanic - Machine Learning from Disaster)
Titanic Dataset Description Overview The data is divided into two groups: - Training set (train.csv): Used to build machine learning models.
2 papers · 1 benchmark
ToM-in-AMC is a novel NLP benchmark, Short for Theory-of-Mind meta-learning Assessment with Movie Characters.
2 papers · 0 benchmarks
TOP is a synthetic dataset for topology optimization generated using Topy.
2 papers · 0 benchmarks
TorWIC (The Toronto Warehouse Incremental Change Dataset)
TorWIC is the dataset discussed in POCD: Probabilistic Object-Level Change Detection and Volumetric Mapping in Semi-Static Scenes.
2 papers · 0 benchmarks
The dataset comprises multiple independent events, where each event contains simulated measurements (essentially 3D points) of particles generated in a collision between proton bunches at the Large Hadron Collider at CERN.
2 papers · 0 benchmarks
Tracking the Trackers is a large-scale analysis of third-party trackers on the World Wide Web.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
Trinity Gesture Dataset includes 23 takes, totalling 244 minutes of motion capture and audio of a male native English speaker producing spontaneous speech on different topics.
2 papers · 2 benchmarks
A benchmark for suppositional reasoning based on the principles of knights and knaves puzzles.
2 papers · 0 benchmarks
The Tsinghua-Daimler Cyclist Benchmark provides a benchmark dataset for cyclist detection.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
TwinViews-13k is a dataset of 13,855 pairs of left-leaning and right-leaning political statements, each pair matched by topic.
2 papers · 0 benchmarks
Twitch-FIFA is video-context, many-speaker dialogue dataset based on live-broadcast soccer game videos and chats from Twitch.tv.
2 papers · 0 benchmarks
The data set contains 2500 manually-stance-labeled tweets, 1250 for each candidate (Joe Biden and Donald Trump).
2 papers · 2 benchmarks
Two-Path Computational Graph (CG) family introduced in "GENNAPE: Towards Generalized Neural Architecture Performance Estimators", accepted to AAAI-23.
2 papers · 0 benchmarks
U-DIADS-Bib is a proprietary dataset developed through the collaboration of computer scientists and humanities at the University of Udine.
2 papers · 1 benchmark
~6 million synthetic depth frames for pose estimation from multiple cameras.
2 papers · 0 benchmarks
A database of several hundred high quality fabric material measurements, provided as carefully calibrated rectified HDR images, together with SVBRDF fits.
2 papers · 0 benchmarks
UCC (Unhealthy Comments Corpus)
The Unhealthy Comments Corpus (UCC) is corpus of 44355 comments intended to assist in research on identifying subtle attributes which contribute to unhealthy conversations online.
2 papers · 0 benchmarks
40,764 images (11,659 protest images and hard negatives) with various annotations of visual attributes and sentiments.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 1 benchmark
This dataset contains 2,000 dial meter images obtained on-site by employees of the Energy Company of Paraná (Copel), which serves more than 4 million consuming units in the Brazilian state of Paraná.
2 papers · 1 benchmark
This dataset contains 2,000 images taken from inside a warehouse of the Energy Company of Paraná (Copel), which directly serves more than 4 million consuming units in the Brazilian state of Paraná.
2 papers · 1 benchmark
The UFPR-Eyeglasses dataset has 1,135 images of both eyes (2,270 cropped images of each eye) from 83 subjects (166 classes).
2 papers · 0 benchmarks
The first ultra-high-definition image demoireing dataset, consisting of 4,500 4K resolution training pairs and 500 standard 4K resolution validation pairs.
2 papers · 1 benchmark
UI5k (Mobile App User Interface Dataset)
This dataset contains 54,987 UI screenshots and the metadata from 7,748 Android applications belonging to 25 application categories Download link: https://www.dropbox.com/sh/kfkhevxykzwputb/AAAhL6ipmOg4zZn4jULmyF0a?dl=0
2 papers · 0 benchmarks
UIT-ViSFD (Vietnamese Aspect-Based Sentiment Analysis Dataset)
UIT-ViSFD is a Vietnamese Smartphone Feedback Dataset as a new benchmark corpus built based on strict annotation schemes for evaluating aspect-based sentiment analysis, consisting of 11,122 human-annotated comments for mobile e-commerce,…
2 papers · 0 benchmarks
UK Biobank participants have generously provided a very wide range of information about their health and well-being since recruitment began in 2006.
2 papers · 1 benchmark
Hundreds of clean and poisoned models per dataset for Tiny-ImageNet, CIFAR10
2 papers · 0 benchmarks
UMLS-43 is a variant of the UMLS knowledge graph that is robust to data leakage through inverse relations.
2 papers · 0 benchmarks
UPFD-POL (User Preference-aware Fake News Detection)
The PolitiFact variant of the UPFD dataset for benchmarking.
2 papers · 1 benchmark
UPenn-GBM (The University of Pennsylvania glioblastoma (UPenn-GBM) cohort)
This collection comprises multi-parametric magnetic resonance imaging (mpMRI) scans for de novo Glioblastoma (GBM) patients from the University of Pennsylvania Health System, coupled with patient demographics, clinical outcome (e.g.,…
2 papers · 0 benchmarks
UQuAD (Urdu Question Answering Dataset)
Large scale machine reading comprehension dataset in Urdu language.
2 papers · 1 benchmark
URBAN-SED is a dataset of 10,000 soundscapes with sound event annotations generated using the scraper library.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
This dataset contains 40,000 URLs of US federal environmental agency websites, along with links to captures in the Internet Archive Wayback Machine for 2016 and 2020 when present.
2 papers · 0 benchmarks
UW Indoor Scenes (UW-IS) Occluded dataset is curated using commodity hardware (Intel RealSense D435) to reflect real world robotics scenarios.
2 papers · 0 benchmarks
UW-IS (UW Indoor Scenes) is a dataset for object recognition in indoor environments comprising scene images from two different environments, namely, a living room and a mock warehouse.
2 papers · 0 benchmarks
The Udacity dataset is mainly composed of video frames taken from urban roads.
2 papers · 1 benchmark
Introduced originally by Xiaohan Yu, Yang Zhao, Yongsheng Gao, Xiaohui Yuan, Shengwu Xiong (2021).
2 papers · 0 benchmarks
The raw data are obtained from an industrial plant for ultra-processed food production.
2 papers · 0 benchmarks
Identify nerve structures in ultrasound images of the neck
2 papers · 0 benchmarks
This dataset contains vibration data recorded on a rotating drive train.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.