Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 61 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 2881–2928 of 12,172
The EARS-WHAM dataset mixes speech from the EARS dataset with real noise recordings from the WHAM!
11 papers · 1 benchmark
the YF-E6 emotion dataset using the 6 basic emotion type as keywords on social video-sharing websites including YouTube and Flickr, leading to a total of 3000 videos.
11 papers · 1 benchmark
FLIP (Fitness Landscape Inference for Proteins)
FLIP includes several benchmark datasets that contain a variety of protein sequences, each with a real-valued label indicating its "fitness" (how well the protein performs some particular function).
11 papers · 0 benchmarks
Fashion 144K is a novel heterogeneous dataset with 144,169 user posts containing diverse image, textual and meta information.
11 papers · 0 benchmarks
FoodX-251 is a dataset of 251 fine-grained classes with 118k training, 12k validation and 28k test images.
11 papers · 1 benchmark
Predicting forest cover type from cartographic variables only (no remotely sensed data).
11 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
11 papers · 0 benchmarks
GSL (Greek Sign Language)
Dataset Description The Greek Sign Language (GSL) is a large-scale RGB+D dataset, suitable for Sign Language Recognition (SLR) and Sign Language Translation (SLT).
11 papers · 1 benchmark
The GTA Indoor Motion dataset (GTA-IM) that emphasizes human-scene interactions in the indoor environments.
11 papers · 2 benchmarks
Four pathologists from Longhua Hospital Shanghai University of Traditional Chinese Medicine provide 600 images of gastric cancer pathology images at size 2048×2048 pixels.
11 papers · 1 benchmark
The human-Related version of the CUHK Avenue dataset, first presented by Morais et al.
11 papers · 1 benchmark
Hilti SLAM Challenge is a dataset for Simultaneous Localization and Mapping (SLAM) algorithms due to sparsity, varying illumination conditions, and dynamic objects.
11 papers · 0 benchmarks
A rich, extensible and efficient environment that contains 45,622 human-designed 3D scenes of visually realistic houses, ranging from single-room studios to multi-storied houses, equipped with a diverse set of fully labeled 3D objects,…
11 papers · 0 benchmarks
A large-scale indoor layout dataset containing 35,357 2D floor plans including 252,550 rooms in total.
11 papers · 0 benchmarks
Housekeep a benchmark to evaluate common sense reasoning in the home for embodied AI.
11 papers · 0 benchmarks
ICVL is a hyperspectral image dataset, collected by "Sparse Recovery of Hyperspectral Signal from Natural RGB Images" The database images were acquired using a Specim PS Kappa DX4 hyperspectral camera and a rotary stage for spatial…
11 papers · 0 benchmarks
A popular dataset for node classification on heterogeneous graphs.
11 papers · 1 benchmark
The ISIC 2018 dataset was published by the International Skin Imaging Collaboration (ISIC) as a large-scale dataset of dermoscopy images.
11 papers · 0 benchmarks
ImageCoDe (Image Retrieval from Contextual Descriptions)
Given 10 minimally contrastive (highly similar) images and a complex description for one of them, the task is to retrieve the correct image.
11 papers · 1 benchmark
ImageNet-X is a set of human annotations pinpointing failure types for the popular ImageNet dataset.
11 papers · 0 benchmarks
JetClass (A Large-Scale Dataset for Deep Learning in Jet Physics)
JetClass is a new large-scale dataset to facilitate deep learning research in particle physics.
11 papers · 1 benchmark
JetNet is a particle cloud dataset, containing gluon, top quark, light quark jets saved in .csv format.
11 papers · 0 benchmarks
Kaggle EyePACS (Kaggle EyePACS. Diabetic Retinopathy Detection Identify signs of diabetic retinopathy in eye images)
Diabetic retinopathy is the leading cause of blindness in the working-age population of the developed world.
11 papers · 1 benchmark
Kobest is a benchmark for Korean language reasoning.
11 papers · 0 benchmarks
A total of 170 videos for training and 30 videos for testing, each of which has 60 frames, amounting to 12,000 paired data.
11 papers · 1 benchmark
LOL-v2-synthetic (From Fidelity to Perceptual Quality: A Semi-Supervised Approach for Low-Light Image Enhancement)
From Fidelity to Perceptual Quality: A Semi-Supervised Approach for Low-Light Image Enhancement
11 papers · 1 benchmark
A large-scale Landmark guided face Parsing dataset (LaPa) for face parsing.
11 papers · 1 benchmark
LandCover.ai (Dataset for Automatic Mapping of Buildings, Woodlands, Water and Roads from Aerial Imagery)
The LandCover.ai (Land Cover from Aerial Imagery) dataset is a dataset for automatic mapping of buildings, woodlands, water and roads from aerial images.
11 papers · 1 benchmark
M4Raw (A multi-contrast, multi-repetition, multi-channel MRI k-space dataset for low-field MRI research)
Recently, low-field magnetic resonance imaging (MRI) has gained renewed interest to promote MRI accessibility and affordability worldwide.
11 papers · 0 benchmarks
The MAGE dataset provides a large set of generated texts using 27 LLMs from seven different groups: OpenAI GPT, LLaMA, GLM130B, FLAN-T5, OPT, BigScience, and EleutherAI.
11 papers · 1 benchmark
MAVEN-ERE is a dataset designed for event relation extraction tasks containing 103,193 event coreference chains, 1,216,217 temporal relations, 57,992 causal relations, and 15,841 subevent relations.
11 papers · 0 benchmarks
MultI-Modal In-Context Instruction Tuning (MIMIC-IT) is a dataset for instruction tuning into multi-modal models, motivated by the Flamingo model's upstream interleaved format pretraining dataset.
11 papers · 0 benchmarks
A data-set which consists of over one million images of physical 3D objects with seven factors of variation, such as object color, shape, size and position.
11 papers · 0 benchmarks
Multi-Modal Reading (MMR) Benchmark includes 550 annotated question-answer pairs across 11 distinct tasks involving texts, fonts, visual elements, bounding boxes, spatial relations, and grounding, with carefully designed evaluation metrics.
11 papers · 1 benchmark
MSD (Million Song Dataset)
The Million Song Dataset is a freely-available collection of audio features and metadata for a million contemporary popular music tracks.
11 papers · 2 benchmarks
Mario AI was a benchmark environment for reinforcement learning.
11 papers · 0 benchmarks
The Matbench test suite v0.1 contains 13 supervised ML tasks from 10 datasets.
11 papers · 0 benchmarks
The Met dataset is a large-scale dataset for Instance-Level Recognition (ILR) in the artwork domain.
11 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
11 papers · 1 benchmark
We scrape data from GooBix, which contains 156 games of 5 × 5 mini crosswords.
11 papers · 0 benchmarks
In this paper, we introduce the MoisesDB dataset for musical source separation.
11 papers · 0 benchmarks
MultiEURLEX is a multilingual dataset for topic classification of legal documents.
11 papers · 0 benchmarks
MULTITQ is a large-scale dataset featuring ample relevant facts and multiple temporal granularities.
11 papers · 1 benchmark
NCBI Datasets is a valuable resource that simplifies the process of gathering data from various NCBI databases.
11 papers · 0 benchmarks
NPHardEval is a dynamic benchmark designed to assess the reasoning abilities of Large Language Models (LLMs) across a broad spectrum of algorithmic questions.
11 papers · 0 benchmarks
The NewSHead dataset contains 369,940 English stories with 932,571 unique URLs, among which there are 359,940 stories for training, 5,000 for validation, and 5,000 for testing, respectively.
11 papers · 1 benchmark
Object HalBench is a benchmark used to evaluate the performance of Language Models, particularly those that are multimodal (i.e., they can process and generate both text and images).
11 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.