Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 146 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 6961–7008 of 12,172
TBBR (Thermal Bridges on Building Rooftops)
The dataset of Thermal Bridges on Building Rooftops (TBBR dataset) consists of annotated combined RGB and thermal drone images with a height map.
2 papers · 2 benchmarks
AI for science has generated a great deal of enthusiasm from both academia and industry.
2 papers · 0 benchmarks
This collection includes datasets from 20 subjects with primary newly diagnosed glioblastoma who were treated with surgery and standard concomitant chemo-radiation therapy (CRT) followed by adjuvant chemotherapy.
2 papers · 0 benchmarks
TCP-CI (Test Case Prioritization in CI Contexts)
This dataset is a benchmark of 25 open-source subjects with 21.5k builds and 3.6k failed builds that enables a fair comparison and evaluation of Test Case Prioritization (TCP) techniques.
2 papers · 0 benchmarks
A new text effects dataset with 141,081 text effect/glyph pairs in total.
2 papers · 0 benchmarks
TEP (Tennessee Eastman Process)
The original paper presented a model of the industrial chemical process named Tennessee Eastman Process and a model-based TEP simulator for data generation.
2 papers · 1 benchmark
THFOOD-50 (Thai Food 50 Image Classification)
Fine-Grained Thai Food Image Classification Datasets THFOOD-50 containing 15,770 images of 50 famous Thai dishes.
2 papers · 0 benchmarks
TI1K Dataset (Thumb Index 1000 Hand & Fingertip Detection Dataset)
Thumb Index 1000 (TI1K) is a dataset of 1000 hand images with the hand bounding box, and thumb and index fingertip positions.
2 papers · 0 benchmarks
TIE (https://github.com/raianand1991/TIE)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
The TII-SSRC-23 dataset offers a comprehensive collection of network traffic patterns, meticulously compiled to support the development and research of Intrusion Detection Systems (IDS).
2 papers · 3 benchmarks
TLDR9+ is a large-scale summarization dataset containing over 9 million training instances extracted from Reddit discussion forum.
2 papers · 1 benchmark
The Text-Music-Dance (TMD) dataset establishes a pioneering benchmark comprising 2,153 text-music-motion pairs.
2 papers · 1 benchmark
TNCR Dataset (Table Net Detection and Classification Dataset)
We present TNCR, a new table dataset with varying image quality collected from free open source websites.
2 papers · 0 benchmarks
A procedurally generated jump'n'run game with control over level similarity.
2 papers · 0 benchmarks
TOMG-Bench (Text-based Open Molecule Generation Benchmark)
In this paper, we propose Text-based Open Molecule Generation Benchmark (TOMG-Bench), the first benchmark to evaluate the open-domain molecule generation capability of LLMs.
2 papers · 1 benchmark
TRN (Toulouse Road Network)
The Toulouse Road Network dataset describes patches of road maps from the city of Toulouse, represented both as spatial graphs G = (A, X) and as grayscale segmentation images.
2 papers · 1 benchmark
TTC (Tatoeba Translation Challenge)
This is a challenge set for machine translation that contains 32G translation units in 2,539 bitexts.
2 papers · 0 benchmarks
TTE-A&O (Travel Time Estimation: Abakan and Omsk)
The dataset includes two parts corresponding to the cities of Abakan (65524 nodes, 340012 edges) and Omsk (231688 nodes, 1149492 edges).
2 papers · 1 benchmark
The dataset has 10.5 hours from a single speaker.
2 papers · 0 benchmarks
TUSC (Tweets from US and Canada)
Tweets from US and Canada (TUSC) is a large dataset of more than 45 million geo-located tweets posted between 2015 and 2021 from US and Canada (TUSC), especially curated for natural language analysis
2 papers · 0 benchmarks
The dataset for this task is the TUT Urban Acoustic Scenes 2018 dataset, consisting of recordings from various acoustic scenes.
2 papers · 1 benchmark
TVIL (Temporal Video Inpainting Localization)
Temporal Video Inpainting Localization Dataset.
2 papers · 0 benchmarks
Touch is an important sensing modality for humans, but it has not yet been incorporated into a multimodal generative language model.
2 papers · 0 benchmarks
TabLeX is a large-scale benchmark dataset comprising table images generated from scientific articles.
2 papers · 0 benchmarks
Talk2Nav is a large-scale dataset with verbal navigation instructions.
2 papers · 0 benchmarks
The TaoDescribe dataset contains 2,129,187 product titles and descriptions in Chinese.
2 papers · 0 benchmarks
This mouse cerebellar atlas can be used for mouse cerebellar morphometry.
2 papers · 0 benchmarks
TeachMyAgent (TA) is a benchmark for Automatic Curriculum Learning (ACL) algorithms leveraging procedural task generation.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
Teeth3DS+ (An Extended Benchmark for Intraoral 3D Scans Analysis)
Intraoral 3D scans analysis is a fundamental aspect of Computer-Aided Dentistry (CAD) systems, playing a crucial role in various dental applications, including teeth segmentation, detection, labeling, and dental landmark identification.
2 papers · 0 benchmarks
README Created by Malireddy Chanakya & Srivenkata N Mounika Somisetty & Malireddy Chaitanya The dataset contains: - 200 short stories - 200 corresponding telegraphic summaries - 50 selected abstractive summaries - 50 selected extractive…
2 papers · 0 benchmarks
TempQA-WD is a benchmark dataset for temporal reasoning designed to encourage research in extending the present approaches to target a more challenging set of complex reasoning tasks.
2 papers · 1 benchmark
It is desirable for detection and classification algorithms to generalize to unfamiliar environments, but suitable benchmarks for quantitatively studying this phenomenon are not yet available.
2 papers · 0 benchmarks
Prototyping complex computer-aided design (CAD) models in modern softwares can be very time-consuming.
2 papers · 0 benchmarks
Text2KGBench is a benchmark to evaluate the capabilities of language models to generate KGs from natural language text guided by an ontology.
2 papers · 0 benchmarks
Tweets and items from psychological scales for sexism detection with counterfactual examples.
2 papers · 0 benchmarks
The Clarity Speech Corpus is a forty speaker British English speech dataset.
2 papers · 0 benchmarks
The Copiale Cipher is a 105 pages manuscript containing all in all around 75 000 characters.
2 papers · 0 benchmarks
The 2048 game task involves training an agent to achieve high scores in the game 2048 (Wikipedia))
2 papers · 1 benchmark
The ULS23 test set contains 725 lesions from 284 patients of the Radboudumc and JBZ hospitals in the Netherlands.
2 papers · 1 benchmark
Large-scale collection of machine learning datasets containing numerical simulations of a wide variety of spatiotemporal physical systems.
2 papers · 0 benchmarks
Involves a crawler to collect data from the Google Play store including the application's metadata and APK files.
2 papers · 0 benchmarks
ThermoHands is the first benchmark dataset specifically designed for egocentric 3D hand pose estimation from thermal images.
2 papers · 0 benchmarks
Data 1: Raw and Unlabeled; 2 million unlabeled replies from 17 Telegram channels.
2 papers · 1 benchmark
A movie ticketing dialog dataset with 23,789 annotated conversations.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.