Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 218 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 10417–10464 of 12,172
This dataset involves a 2D or 3D agent moving from a start to goal pose while interacting with nearby objects.
1 paper · 0 benchmarks
About Dataset The File contains 3D point cloud data of a Fabricate plant with 10 sequences.
1 paper · 0 benchmarks
Overview: This collection contains three synthetic datasets produced by gpt-4o-mini for sentiment analysis and PDT (Product Desirability Toolkit) testing.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Synthetic dataset comprising three different environments for multi-camera dynamic novel view synthesis for soccer.
1 paper · 0 benchmarks
Synthetic Speech Attribution Dataset.
1 paper · 0 benchmarks
Synthetic visual inspection data of structural elements in bridges.
1 paper · 0 benchmarks
Each file contains a specific dataset described in the paper "On Automatic Parsing of Log Records".
1 paper · 0 benchmarks
Generated using the script below: https://github.com/zenineasa/MasterThesis/blob/main/Code/dataGenerator.py
1 paper · 0 benchmarks
We used the following procedure.
1 paper · 0 benchmarks
The dataset contains naive and stylized sketches for a chair category of the ShapeNetCore dataset.
1 paper · 0 benchmarks
SyntheticFur is a dataset for neural rendering.
1 paper · 0 benchmarks
T2 Guiding is a dataset of 1000 images, each with six image labels.
1 paper · 0 benchmarks
Table-ACM12K (TACM12K) is a relational table dataset derived from the ACM heterogeneous graph dataset.
1 paper · 1 benchmark
TACO-BAAI (Topics in Algorithmic Code generation dataset)
TACO (Topics in Algorithmic Code generation dataset) is a dataset focused on algorithmic code generation, designed to provide a more challenging training dataset and evaluation benchmark for the code generation model field.
1 paper · 1 benchmark
TADAC (Text Annotated Distortion, Appearance and Content Dataset)
We have developed a systematic method for constructing large text annotated image databases designed for exploiting vision-language modeling for image quality assessment and present the Text Annotated Distortion, Appearance and Content…
1 paper · 0 benchmarks
TAMPAR is a real-world dataset of parcel photos for tampering detection with annotations in COCO format.
1 paper · 0 benchmarks
Our dataset augments the TAO dataset with amodal bounding box annotations for fully invisible, out-of-frame, and occluded objects.
1 paper · 0 benchmarks
TARA is a dataset for tool-augmented reward modeling, which includes comprehensive comparison data of human preferences and detailed tool invocation processes.
1 paper · 0 benchmarks
TASTEset Recipe Dataset and Food Entities Recognition is a dataset for Named Entity Recognition (NER) which consists of 700 recipes with more than 13,000 entities to extract.
1 paper · 0 benchmarks
TAT (Taiwanese Across Taiwan)
Taiwanese Across Taiwan (TAT) corpus is a Large-Scale database of Native Taiwanese Article/Reading Speech collected across Taiwan.
1 paper · 1 benchmark
The dataset for this task is TAU Audio-Visual Urban Scenes 2021.
1 paper · 0 benchmarks
TB-Places is a data set of garden images for testing algorithms for visual place recognition.
1 paper · 0 benchmarks
TBBR Raw (Hyperspectral (RGB + Thermal) drone images of Karlsruhe, Germany)
This dataset contains the raw images for the dataset of Thermal Bridges on Building Rooftops (TBBR) dataset.
1 paper · 0 benchmarks
TBCOV is a large-scale Twitter dataset comprising more than two billion multilingual tweets related to the COVID-19 pandemic collected worldwide over a continuous period of more than one year.
1 paper · 0 benchmarks
TCB-DS (Toxigenic Cyanobacteria Dataset)
The TCB-DS dataset is a specialized collection of microscopic images focusing on the automatic recognition of cyanobacteria genera.
1 paper · 0 benchmarks
A dataset of 663 deidentified computed tomography (CT) scans acquired in routine clinical practice and with both segmentations taken from clinical practice and segmentations.
1 paper · 0 benchmarks
TCLD (Typhoon Center Location Dataset)
TCLD (Typhoon Center Location Dataset) is a brand new typhoon center location dataset for deep learning research.
1 paper · 0 benchmarks
TCR-CMV (T cell repertoires labelled by CMV serostatus)
Adaptive Biotechnologies' dataset of sequenced T cell repertoires labelled by patient age, HLA type, and CMV serostatus Source: Immunosequencing identifies signatures of cytomegalovirus exposure history and HLA-mediated effects on the…
1 paper · 0 benchmarks
TCR-pMHC (10x Genomics T cell receptor peptide-MHC pairs)
10x Genomics dataset of sequenced TCRs barcoded by a panel of pMHCs (arranged on a dextramer) Source: A new way of exploring immunity: linking highly multiplexed antigen recognition to immune repertoire and phenotype
1 paper · 0 benchmarks
This dataset is a collection of paired wireless signal data and corresponding image ground truth specifically designed for underground object sensing and image reconstruction.
1 paper · 0 benchmarks
TDMD contains eight reference DCM objects with six typical distortions.
1 paper · 0 benchmarks
The TED VCR Video Retrieval Dataset is a multimodal collection derived from publicly available TED Talks.
1 paper · 0 benchmarks
TEM image dataset containing four nanowire morphologies of bio-derived protein nanowires and synthetic peptide nanowires.
1 paper · 0 benchmarks
TERRA-REF (TERRA-REF, An open reference data set from high resolution genomics, phenomics, and imaging sensors)
The ARPA-E funded TERRA-REF project is generating open-access reference datasets for the study of plant sensing, genomics, and phenomics.
1 paper · 0 benchmarks
A collection of photographic and synthetic images intended for analysis of image processing techniques and quality assessment of displays.
1 paper · 0 benchmarks
TF1-EN-3M: Three Million Synthetic Moral Fables for Open Language Models TF1-EN-3M is a large-scale synthetic dataset of 3,000,000 English-language moral fables, generated by instruction-tuned language models with no more than 8 billion…
1 paper · 0 benchmarks
TFRD (Temperature Field Reconstruction Dataset)
TFRD is a dataset to evaluate machine learning modelling methods for theTemperature field reconstruction of heat source systems (TFR-HSS).
1 paper · 0 benchmarks
TFVulFix (TensorFlow Vulnerability Fixes)
TFVulFix is a dataset containing commits from TensorFlow, which is a well-known deep learning library.
1 paper · 0 benchmarks
The dataset contains more than 100k code patch pairs extracted from open source projects on GitHub.
1 paper · 1 benchmark
TGB (Temporal Graph Benchmark)
TGB is a collection of challenging and diverse benchmark datasets for realistic, reproducible, and robust machine learning evaluation on temporal graphs.
1 paper · 0 benchmarks
TGRDB (Tour-Guide Robot Dataset and Benchmark)
1.
1 paper · 0 benchmarks
THEOStereo is a dataset providing synthetic stereo image pairs and their corresponding scene depth and will be published along with [1].
1 paper · 0 benchmarks
THGP (Temporal Hands Guns and Phones Dataset)
Temporal Hands Guns and Phones (THGP) dataset, is a collection of 5960 video frames (5000 for training and 960 for testing).
1 paper · 0 benchmarks
THRED (Two-Hop Relation Extraction Dataset)
This is two-hop relation extraction dataset derived from WikiHop dataset [1].
1 paper · 0 benchmarks
THRawS (Thermal Hotspots in Raw Sentinel-2 data)
THRawS is a new dataset of raw Sentinel-2 (S-2) satellite data containing warm temperature hotspots such as wildfires and volcanic eruptions from around the world.
1 paper · 0 benchmarks
THU-FVFDT (Tsinghua University Finger Vein and Finger Dorsal Texture Database)
THU-FVFDT is a dataset containing raw finger vein and finger dorsal texture images of 220 different subjects.
1 paper · 0 benchmarks
THÖR is a dataset with human motion trajectory and eye gaze data collected in an indoor environment with accurate ground truth for position, head orientation, gaze direction, social grouping, obstacles map and goal coordinates.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.