Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 220 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 10513–10560 of 12,172
TeleSim (TeleSim: A Network-Aware Testbed and Benchmark Dataset for Telerobotic Applications)
TeleSim is a network-aware hardware-in-the-loop dataset designed to evaluate the performance of telerobotic systems under varying network conditions.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
TempWikiBio is a new data-to-text generation dataset containing more than 4 millions of chronologically ordered revisions of biographical articles from English Wikipedia, each paired with structured personal profiles.
1 paper · 0 benchmarks
Data URL: https://data.mendeley.com/datasets/9k892pzkfx/1
1 paper · 0 benchmarks
This is a comprehensive test database of scenes that treat different light setups in conjunction with diverse materials.
1 paper · 0 benchmarks
Green family of datasets for emergent communications on relations.
1 paper · 0 benchmarks
A large dataset of natural language descriptions for physical 3D objects in the ShapeNet dataset.
1 paper · 0 benchmarks
We introduce TextAtlas5M, a dataset specifically designed for training and evaluating multimodal generation models on dense-text image generation.
1 paper · 0 benchmarks
Text present in images are not merely strings, they provide useful cues about the image.
1 paper · 0 benchmarks
TextWorld KG is a dynamic Knowledge Graph (KG) extraction dataset.
1 paper · 0 benchmarks
Extends the COCO-text [Veit et al.
1 paper · 0 benchmarks
Texygen is a benchmarking platform to support research on open-domain text generation models.
1 paper · 0 benchmarks
The Benchmark is a collection of datasets for Monocular Height Estimation.
1 paper · 0 benchmarks
The Berka dataset is a collection of financial information from a Czech bank.
1 paper · 0 benchmarks
Content This dataset contains all utterances of two episodes of South Park (Latin American voices) and two episodes of Archer (Spanish voices).
1 paper · 0 benchmarks
The ComMA Dataset v0.2 is a multilingual dataset annotated with a hierarchical, fine-grained tagset marking different types of aggression and the "context" in which they occur.
1 paper · 0 benchmarks
Using the Experience-Sampling Method (ESM), participants are asked to report TV consumption multiple times each day for a five week period.
1 paper · 0 benchmarks
We present the SourceData-NLP dataset produced through the routine curation of papers during the publication process.
1 paper · 1 benchmark
Hyperspectral Imaging, employed in satellites for space remote sensing, like HYPSO-1, faces constraints due to few labeled data sets, affecting the training of AI models demanding these ground-truth annotations.
1 paper · 0 benchmarks
The Mafia Dataset was created to model the behavior of deceptive actors in the context of the Mafia game, as described in the paper “Putting the Con in Context: Identifying Deceptive Actors in the Game of Mafia”.
1 paper · 0 benchmarks
The RBO dataset of articulated objects and interactions is a collection of 358 RGB-D video sequences (67:18 minutes) of humans manipulating 14 articulated objects under varying conditions (light, perspective, background, interaction).
1 paper · 0 benchmarks
The Reddit Climate Change Dataset is a dataset of 620K Reddit posts and 4.6M comments - all mentions of the terms "climate" and "change" until 2022-09-01 across the entire Reddit social network.
1 paper · 0 benchmarks
This includes all data from the ACM IMC 2018 paper "The Rise of Certificate Transparency and Its Implications on the Internet Ecosystem".
1 paper · 0 benchmarks
In this paper, we introduce a victim dataset for the RoboCup Rescue competitions.
1 paper · 0 benchmarks
Photorealistic indoor dataset designed to enable the application of deep learning techniques to a wide variety of robotic vision problems.
1 paper · 0 benchmarks
The SWC is a corpus of aligned Spoken Wikipedia articles from the English, German, and Dutch Wikipedia.
1 paper · 1 benchmark
The TREC Fair Ranking track evaluates systems according to how well they fairly rank documents.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The ULS23 training dataset contains 38,693 diverse lesions from chest-abdomen-pelvis CT examinations.
1 paper · 0 benchmarks
We present a new annotated corpus of written learner English, derived from essays submitted to the learning platform Write & Improve (W&I).
1 paper · 0 benchmarks
High-resolution thermal infrared face database with extensive manual annotations, introduced by Kopaczka et al, 2018.
1 paper · 0 benchmarks
The database was acquired using a thermographic camera TESTO 880-3.
1 paper · 0 benchmarks
ThermalWORLD-C is an evaluation set that consists of algorithmically generated corruptions applied to the ThermalWORLD test-set, and especially to both the visible and the thermal data.
1 paper · 0 benchmarks
Dataset of paired thermal and RGB images comprising ten diverse scenes—six indoor and four outdoor scenes— for 3D scene reconstruction and novel view synthesis (e.g.
1 paper · 0 benchmarks
This is not a Dataset (This is not a Dataset: A Large Negation Benchmark to Challenge Large Language Models)
We introduce a large semi-automatically generated dataset of ~400,000 descriptive sentences about commonsense knowledge that can be true or false in which negation is present in about 2/3 of the corpus in different forms that we use to…
1 paper · 1 benchmark
Thorsten-Voice (Thorsten-21.02-neutral) is a neutrally spoken voice dataset recorded by Thorsten Müller, audio optimized by Dominik Kreutz and licenced under CC0 to provide it for anybody without any financial or licence struggle.
1 paper · 1 benchmark
10000 instances of three-view numerical data set with 4 clusters and 2 feature components are considered.
1 paper · 0 benchmarks
Thunder-NUBench (Negation Understanding Benchmark) is a benchmark specifically designed to evaluate large language models’ (LLMs) sentence-level understanding of negation.
1 paper · 0 benchmarks
TiROD (Tiny Robotics Object Detection)
Dataset to benchmark Continual Learning for Object Detection in a Tiny Robotics settings.
1 paper · 1 benchmark
TikTok Comments is a domain specific lexicon based on TikTok comments dataset.
1 paper · 0 benchmarks
Includes considerable roll and pitch camera motion.
1 paper · 0 benchmarks
The TimberVision dataset consists of more than 2k annotated RGB images and contains a total of 51k trunk components including cut and lateral surfaces, thereby surpassing any existing dataset in this domain in terms of both quantity and…
1 paper · 0 benchmarks
TimeGraph (TimeGraph: Synthetic Benchmark Datasets for Robust Time-Series Causal Discovery)
TimeGraph is a comprehensive suite of synthetic datasets designed to benchmark causal discovery algorithms on time-series data.
1 paper · 0 benchmarks
Question answering over temporal knowledge graphs (TKGs) is crucial for understanding evolving facts and relationships, yet its development is hindered by limited datasets and difficulties in generating custom QA pairs.
1 paper · 0 benchmarks
Tinto (Tinto: Multisensor Benchmark for 3D Hyperspectral Point Cloud Segmentation in the Geosciences)
The increasing use of deep learning techniques has reduced interpretation time and, ideally, reduced interpreter bias by automatically deriving geological maps from digital outcrop models.
1 paper · 0 benchmarks
Tiny ImageNet-A is a subset of the Tiny ImageNet test set consisting of 3,374 images comprising real-world, unmodified, and naturally occurring examples that are misclassified by ResNet-18.
1 paper · 0 benchmarks
TinyChirp dataset for model training, validation and testing
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.