Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 221 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 10561–10608 of 12,172
The Tiny Shakespeare corpus is a dataset that contains 40,000 lines of Shakespeare from a variety of his plays.
1 paper · 0 benchmarks
TinySocial is a dataset to enable research on Social Visual Question Answering.
1 paper · 0 benchmarks
TinyVIRAT-v2 is a benchmark dataset for recognizing real-world low-resolution activities present in videos.
1 paper · 0 benchmarks
The data consists of a set of 3 task types and 4 question types, creating 12 total scenarios.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
A dataset made of 3D image data and their embeddings to test TomoSAM
1 paper · 0 benchmarks
Tool Database for image-set clustering This database was generated to evaluate a robotic application dealing with image-set clustering.
1 paper · 0 benchmarks
The dataset is organized as follows.
1 paper · 0 benchmarks
A set of Monte Carlo simulated events, for the evaluation of top quarks' (and their child particles') momentum reconstruction, produced using the HEPData4ML package [1].
1 paper · 0 benchmarks
A prevalent use case of topic models is that of topic discovery.
1 paper · 1 benchmark
The Touché23-ValueEval Dataset is a collection of arguments used for identifying human values behind those arguments.
1 paper · 0 benchmarks
Toulouse Vanishing Points Dataset is a public photographs database of Manhattan scenes taken with an iPad Air 1.
1 paper · 0 benchmarks
6000 French user reviews from three applications on Google Play (Garmin Connect, Huawei Health, Samsung Health) are labelled manually.
1 paper · 0 benchmarks
Toxicity (Toxicity - UCI Machine Learning Repository)
Classification dataset of 171 molecules according to its toxicity from the UCI Machine Learning Repository.
1 paper · 0 benchmarks
A new high accuracy Turkish morphology dataset.
1 paper · 0 benchmarks
TrUMAn (Trope Understanding in Movies and Animations)
Trope Understanding in Movies and Animations (TrUMAn) is a dataset intending to evaluate and develop learning systems beyond visual signals.
1 paper · 0 benchmarks
TraVLR is a synthetic dataset comprising four visio-linguistic reasoning tasks.
1 paper · 0 benchmarks
This dataset consists of 2,192 high-quality traditional Chinese landscape paintings (中国山水画).
1 paper · 0 benchmarks
This data set is being released to support the spam and context-specific spam detection tasks on Twitter data.
1 paper · 3 benchmarks
This dataset includes four real-world sub-datasets about traffic demand.
1 paper · 0 benchmarks
Trailers12k is a movie trailer dataset comprised of 12,000 titles associated to ten genres.
1 paper · 0 benchmarks
This repository contains data for the NeurIPS conference paper titled "Harnessing Machine Learning for Single-Shot Measurement of Free Electron Laser Pulse Power".
1 paper · 0 benchmarks
Data and experiments for motion-based extrinsic calibration using [trajectorycalibration].
1 paper · 0 benchmarks
The dataset contains procedurally generated images of transparent vessels containing liquid and objects .
1 paper · 1 benchmark
Translated SNLI Dataset in Marathi A translated version of the SNLI dataset in Marathi, designed for Semantic Textual Similarity (STS) tasks.
1 paper · 1 benchmark
533 parallel examples sampled from TACRED, translated into Russian and Korean (and 3 additional examples in Russian), accompanied with tranlsation of a list of trigger words collected for the different relations.
1 paper · 0 benchmarks
Dataset of the RANS simulations over a 2D RAE2822 Airfoil at different Mach and Angle of Attack.
1 paper · 0 benchmarks
Trefoil100 is a dataset to test knot untangling algorithms, i.e., highly-tangled configurations that can be difficult to smooth out into a canonical knot embedding.
1 paper · 0 benchmarks
Source: Reconstructing lineage hierarchies of the distal lung epithelium using single-cell RNA-seq
1 paper · 1 benchmark
Twitter dataset related to flood events onsets in Thailand and Nepal, focused on September 26/27, 2022, June 16/17 2021 and July 01/02 2021.
1 paper · 0 benchmarks
The Trilemma of Truth is a multiclass probing dataset for evaluating the veracity-tracking mechanism of large language models.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The dataset contains 1,200 trained ViT-B-16 models, trained on ImageNet.
1 paper · 0 benchmarks
Data used in the work ”Unifying the design space and optimizing linear and nonlinear truss metamaterials by generative modeling”.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
includes train, test and validation set
1 paper · 0 benchmarks
TruthGen is a dataset of generated true and false statements, intended for research on truthfulness in reward models and language models, specifically in contexts where political bias is undesirable.
1 paper · 0 benchmarks
Tsinghua Dogs is a fine-grained classification dataset for dogs, over 65% of whose images are collected from people's real life.
1 paper · 0 benchmarks
TuGebic (A Turkish-German Bilingual Code-Switching Corpus)
TuGebic is a corpus of recordings of spontaneous speech samples from Turkish-German bilinguals, and the compilation of a corpus called TuGebic.
1 paper · 0 benchmarks
TuPyE, an enhanced iteration of TuPy, encompasses a compilation of 43,668 meticulously annotated documents specifically selected for the purpose of hate speech detection within diverse social network contexts.
1 paper · 0 benchmarks
The TupleInf Open IE dataset contains Open IE tuples extracted from 263K sentences that were used by the solver in “Answering Complex Questions Using Open Information Extraction” (referred as Tuple KB, T).
1 paper · 0 benchmarks
Turath-150K is a database of images of the Arab world that reflect objects, activities, and scenarios commonly found there.
1 paper · 0 benchmarks
Turbulence is a new benchmark for systematically evaluating the correctness and robustness of instruction-tuned large language models (LLMs) for code generation.
1 paper · 1 benchmark
--- license: mit language: - en tags: - urbanMapping - 3d - pointCloud - LiDAR prettyname: Turin3D sizecategories: - 10M dataset.tar.gz tar -xvzf dataset.tar.gz Class Taxonomy The dataset utilizes a taxonomy of 6 semantic classes: 0.
1 paper · 0 benchmarks
TurkQA consists of a selection of sentences from English Wikipedia articles, with questions and answers crowdsourced from workers on Amazon Mechanical Turk.
1 paper · 0 benchmarks
we have prepared a dataset using publicly available TED Talks transcripts [27] and selected the Turkish corpus.
1 paper · 0 benchmarks
This dataset is a record of the active learning data collected from interacting with PersonaGPT to fine-tune its actions toward turn-level goals, which are text descriptions of decoding goals for each response in a conversation.
1 paper · 0 benchmarks
TwBNT is the bot detection benchmark with automatic troll annotations.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.