Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 17 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 769–816 of 12,172
The How2 dataset contains 13,500 videos, or 300 hours of speech, and is split into 185,187 training, 2022 development (dev), and 2361 test utterances.
84 papers · 2 benchmarks
MiniF2F is a dataset of formal Olympiad-level mathematics problems statements intended to provide a unified cross-system benchmark for neural theorem proving.
84 papers · 2 benchmarks
MusicCaps is a dataset composed of 5.5k music-text pairs, with rich text descriptions provided by human experts.
84 papers · 1 benchmark
The PROMISE12 dataset was made available for the MICCAI 2012 prostate segmentation challenge.
84 papers · 2 benchmarks
TweetEval introduces an evaluation framework consisting of seven heterogeneous Twitter-specific classification tasks.
84 papers · 1 benchmark
VITON-HD (High-Resolution VITON-Zalando Dataset)
VITON-HD dataset is a dataset for high-resolution (i.e., 1024x768) virtual try-on of clothing items.
84 papers · 2 benchmarks
WikiBio (Wikipedia Biography Dataset)
This dataset gathers 728,321 biographies from English Wikipedia.
84 papers · 1 benchmark
GovReport is a dataset for long document summarization, with significantly longer documents and summaries.
83 papers · 2 benchmarks
NLVR (Natural Language Visual Reasoningnatural language for visual reasoning)
NLVR contains 92,244 pairs of human-written English sentences grounded in synthetic images.
83 papers · 3 benchmarks
A large dataset of human hand images (dorsal and palmar sides) with detailed ground-truth information for gender recognition and biometric identification.
82 papers · 0 benchmarks
ABO (Amazon Berkeley Objects)
ABO is a large-scale dataset designed for material prediction and multi-view retrieval experiments.
82 papers · 0 benchmarks
DrawBench is a comprehensive and challenging benchmark for text-to-image models, introduced by the Imagen research team.
82 papers · 1 benchmark
MathVerse is an innovative benchmark specifically designed to rigorously evaluate the capabilities of Multi-modal Large Language Models (MLLMs) in interpreting and reasoning with visual information in mathematical problems.
82 papers · 0 benchmarks
Our task is to localize and provide a pixel-level mask of an object on all video frames given a language referring expression obtained either by looking at the first frame only or the full video.
82 papers · 1 benchmark
CoMA contains 17,794 meshes of the human face in various expressions Source: DEMEA: Deep Mesh Autoencoders for Non-Rigidly Deforming Objects Image Source: https://coma.is.tue.mpg.de/
81 papers · 1 benchmark
ChestX-ray8 is a medical imaging dataset which comprises 108,948 frontal-view X-ray images of 32,717 (collected from the year of 1992 to 2015) unique patients with the text-mined eight common disease labels, mined from the text…
81 papers · 0 benchmarks
Douban (Douban Conversation Corpus)
We release Douban Conversation Corpus, comprising a training data set, a development set and a test set for retrieval based chatbot.
81 papers · 4 benchmarks
The INTERACTION dataset contains naturalistic motions of various traffic participants in a variety of highly interactive driving scenarios from different countries.
81 papers · 1 benchmark
LoveDA (Remote Sensing Land-Cover Dataset for Domain Adaptive Semantic Segmentation)
1.
81 papers · 1 benchmark
The MetaQA dataset consists of a movie ontology derived from the WikiMovies Dataset and three sets of question-answer pairs written in natural language: 1-hop, 2-hop, and 3-hop queries.
81 papers · 1 benchmark
RaFD (Radboud Faces Database)
The Radboud Faces Database (RaFD) is a set of pictures of 67 models (both adult and children, males and females) displaying 8 emotional expressions.
81 papers · 2 benchmarks
The XSTest dataset is a test suite designed to identify exaggerated safety behaviors in large language models.
81 papers · 0 benchmarks
iSAID contains 655,451 object instances for 15 categories across 2,806 high-resolution images.
81 papers · 4 benchmarks
The AMiner Dataset is a collection of different relational datasets.
80 papers · 1 benchmark
LFSD (Light Field Saliency Database)
The Light Field Saliency Database (LFSD) contains 100 light fields with 360×360 spatial resolution.
80 papers · 1 benchmark
Oulu-CASIA (Oulu-CASIA NIR&VIS facial expression database)
The Oulu-CASIA NIR&VIS facial expression database consists of six expressions (surprise, happiness, sadness, anger, fear and disgust) from 80 people between 23 and 58 years old.
80 papers · 4 benchmarks
Places-LT has an imbalanced training set with 62,500 images for 365 classes from Places-2.
80 papers · 1 benchmark
The ProofWriter dataset contains many small rulebases of facts and rules, expressed in English.
80 papers · 0 benchmarks
RealNews is a large corpus of news articles from Common Crawl.
80 papers · 0 benchmarks
The ReferIt dataset contains 130,525 expressions for referring to 96,654 objects in 19,894 images of natural scenes.
80 papers · 0 benchmarks
VQG (Visual Question Generation)
VQG is a collection of datasets for visual question generation.
80 papers · 1 benchmark
Volleyball is a video action recognition dataset.
80 papers · 3 benchmarks
CoNLL-2014 will continue the CoNLL tradition of having a high profile shared task in natural language processing.
79 papers · 0 benchmarks
MOSES (Molecular sets (MOSES))
The set is based on the ZINC Clean Leads collection.
79 papers · 0 benchmarks
VeRi-776 is a vehicle re-identification dataset which contains 49,357 images of 776 vehicles from 20 cameras.
79 papers · 1 benchmark
WIT (Wikipedia-based Image Text)
Wikipedia-based Image Text (WIT) Dataset is a large multimodal multilingual dataset.
79 papers · 1 benchmark
WikiTableQuestions is a question answering dataset over semi-structured tables.
79 papers · 2 benchmarks
MPIIGaze is a dataset for appearance-based gaze estimation in the wild.
78 papers · 2 benchmarks
OPUS-100 is an English-centric multilingual corpus covering 100 languages.
78 papers · 0 benchmarks
OPV2V is a large-scale open simulated dataset for Vehicle-to-Vehicle perception.
78 papers · 2 benchmarks
Over a period of three years (2009 - 2011) the daily news and weather forecast airings of the German public tv-station PHOENIX featuring sign language interpretation have been recorded and the weather forecasts of a subset of 386 editions…
78 papers · 0 benchmarks
Reddit-5K is a relational dataset extracted from Reddit.
78 papers · 1 benchmark
RadGraph (RadGraph: Extracting Clinical Entities and Relations from Radiology Reports)
RadGraph is a dataset of entities and relations in radiology reports based on our novel information extraction schema, consisting of 600 reports with 30K radiologist annotations and 221K reports with 10.5M automatically generated…
78 papers · 0 benchmarks
CoNaLa (CMU CoNaLa, the Code/Natural Language Challenge)
The CMU CoNaLa, the Code/Natural Language Challenge dataset is a joint project from the Carnegie Mellon University NeuLab and Strudel labs.
77 papers · 1 benchmark
Few-NERD is a large-scale, fine-grained manually annotated named entity recognition dataset, which contains 8 coarse-grained types, 66 fine-grained types, 188,200 sentences, 491,711 entities, and 4,601,223 tokens.
77 papers · 3 benchmarks
Kubric is a data generation pipeline for creating semi-realistic synthetic multi-object videos with rich annotations such as instance segmentation masks, depth maps, and optical flow.
77 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.