12,172 datasets listed, ordered by the archive's paper count. Page 10 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
FSD50K (Freesound Database 50K)
Freesound Dataset 50k (or FSD50K for short) is an open dataset of human-labeled sound events containing 51,197 Freesound clips unequally distributed in 200 classes drawn from the AudioSet Ontology.
155 papers · 2 benchmarks
MIND (MIcrosoft News Dataset)
MIcrosoft News Dataset (MIND) is a large-scale dataset for news recommendation research.
155 papers · 0 benchmarks
MTEB (Massive Text Embedding Benchmark)
MTEB is a benchmark that spans 8 embedding tasks covering a total of 56 datasets and 112 languages.
155 papers · 6 benchmarks
ObjectNet is a test set of images collected directly using crowd-sourcing.
155 papers · 4 benchmarks
The See-in-the-Dark (SID) dataset contains 5094 raw short-exposure images, each with a corresponding long-exposure reference image.
155 papers · 3 benchmarks
A-OKVQA is crowdsourced visual question answering dataset composed of a diverse set of about 25K questions requiring a broad base of commonsense and world knowledge to answer.
154 papers · 1 benchmark
AFLW (Annotated Facial Landmarks in the Wild)
The Annotated Facial Landmarks in the Wild (AFLW) is a large-scale collection of annotated face images gathered from Flickr, exhibiting a large variety in appearance (e.g., pose, expression, ethnicity, age, gender) as well as general…
154 papers · 11 benchmarks
AFW (Annotated Faces in the Wild)
AFW (Annotated Faces in the Wild) is a face detection dataset that contains 205 images with 468 faces.
154 papers · 1 benchmark
In-Shop (In-shop Clothes Retrieval Benchmark)
In-shop Clothes Retrieval Benchmark evaluates the performance of in-shop Clothes Retrieval.
154 papers · 2 benchmarks
The NCBI Disease corpus consists of 793 PubMed abstracts, which are separated into training (593), development (100) and test (100) subsets.
154 papers · 3 benchmarks
Conceptual 12M (CC12M) is a dataset with 12 million image-text pairs specifically meant to be used for vision-and-language pre-training.
153 papers · 0 benchmarks
PATTERN is a node classification tasks generated with Stochastic Block Models, which is widely used to model communities in social networks by modulating the intra- and extra-communities connections, thereby controlling the difficulty of…
153 papers · 1 benchmark
dSprites (Disentanglement testing Sprites dataset)
dSprites is a dataset of 2D shapes procedurally generated from 6 ground truth independent latent factors.
153 papers · 0 benchmarks
The Long-tailed Version of CIFAR100 Source: Learning Imbalanced Datasets with Label-Distribution-Aware Margin Loss
152 papers · 0 benchmarks
The MegaDepth dataset is a dataset for single-view depth prediction that includes 196 different locations reconstructed from COLMAP SfM/MVS.
152 papers · 0 benchmarks
NAS-Bench-101 is the first public architecture dataset for NAS research.
152 papers · 1 benchmark
OLID (Offensive Language Identification Dataset)
The OLID is a hierarchical dataset to identify the type and the target of offensive texts in social media.
152 papers · 1 benchmark
ROCStories is a collection of commonsense short stories.
152 papers · 2 benchmarks
A large corpus of 81.1M English-language academic papers spanning many academic disciplines.
152 papers · 1 benchmark
Video-MME stands for Video Multi-Modal Evaluation.
152 papers · 2 benchmarks
FCE (First Certificate in English)
The Cambridge Learner Corpus First Certificate in English (CLC FCE) dataset consists of short texts, written by learners of English as an additional language in response to exam prompts eliciting free-text answers and assessing mastery of…
151 papers · 1 benchmark
SQuAD (Stanford Question Answering Dataset)
The Stanford Question Answering Dataset (SQuAD) is a collection of question-answer pairs derived from Wikipedia articles.
151 papers · 12 benchmarks
VRD (Visual Relationship Detection dataset)
The Visual Relationship Dataset (VRD) contains 4000 images for training and 1000 for testing annotated with visual relationships.
151 papers · 5 benchmarks
ASDiv (Academia Sinica Diverse MWP Dataset)
We present ASDiv (Academia Sinica Diverse MWP Dataset), a diverse (in terms of both language patterns and problem types) English math word problem (MWP) corpus for evaluating the capability of various MWP solvers.
150 papers · 1 benchmark
The MMLU-Pro dataset is an enhanced version of the Massive Multitask Language Understanding (MMLU) benchmark.
150 papers · 1 benchmark
REDDIT-BINARY consists of graphs corresponding to online discussions on Reddit.
150 papers · 1 benchmark
MOT16 (Multiple Object Tracking 2016)
The MOT16 dataset is a dataset for multiple object tracking.
149 papers · 2 benchmarks
The WebNLG corpus comprises of sets of triplets describing facts (entities and relations between them) and the corresponding facts in form of natural language text.
149 papers · 17 benchmarks
DISFA (Denver Intensity of Spontaneous Facial Action)
The Denver Intensity of Spontaneous Facial Action (DISFA) dataset consists of 27 videos of 4844 frames each, with 130,788 images in total.
148 papers · 3 benchmarks
The Kinetics-600 is a large-scale action recognition dataset which consists of around 480K videos from 600 action categories.
148 papers · 3 benchmarks
SCAN (Simplified versions of the CommAI Navigation tasks)
SCAN is a dataset for grounded navigation which consists of a set of simple compositional navigation commands paired with the corresponding action sequences.
148 papers · 0 benchmarks
TyDiQA (Typologically Diverse Question Answering)
TyDi QA is a question answering dataset covering 11 typologically diverse languages with 200K question-answer pairs.
148 papers · 0 benchmarks
The 2D-3D-S dataset provides a variety of mutually registered modalities from 2D, 2.5D and 3D domains, with instance-level semantic and geometric annotations.
147 papers · 6 benchmarks
CrowS-Pairs has 1508 examples that cover stereotypes dealing with nine types of bias, like race, religion, and age.
147 papers · 1 benchmark
A new open-vocabulary language modelling benchmark derived from books.
147 papers · 1 benchmark
SBD (Semantic Boundaries Dataset)
The Semantic Boundaries Dataset (SBD) is a dataset for predicting pixels on the boundary of the object (as opposed to the inside of the object with semantic segmentation).
147 papers · 2 benchmarks
Taskonomy provides a large and high-quality dataset of varied indoor scenes.
147 papers · 2 benchmarks
Urban Sound 8K is an audio dataset that contains 8732 labeled sound excerpts (<=4s) of urban sounds from 10 classes: airconditioner, carhorn, childrenplaying, dogbark, drilling, engingeidling, gunshot, jackhammer, siren, and streetmusic.
147 papers · 1 benchmark
The YouTube-8M dataset is a large scale video dataset, which includes more than 7 million videos with 4716 classes labeled by the annotation system.
147 papers · 2 benchmarks
aPY (Attribute Pascal and Yahoo)
aPY is a coarse-grained dataset composed of 15339 images from 3 broad categories (animals, objects and vehicles), further divided into a total of 32 subcategories (aeroplane, …, zebra).
147 papers · 5 benchmarks
The ActivityNet-QA dataset contains 58,000 human-annotated QA pairs on 5,800 videos derived from the popular ActivityNet dataset.
146 papers · 2 benchmarks
MMStar is an elite vision-indispensable multi-modal benchmark comprising 1,500 meticulously selected samples.
146 papers · 0 benchmarks
STARE (Structured Analysis of the Retina)
The STARE (Structured Analysis of the Retina) dataset is a dataset for retinal vessel segmentation.
146 papers · 6 benchmarks
The SumMe dataset is a video summarization dataset consisting of 25 videos, each annotated with at least 15 human summaries (390 in total).
146 papers · 3 benchmarks
The TVQA dataset is a large-scale video dataset for video question answering.
146 papers · 3 benchmarks
A new dataset with abstractive dialogue summaries.
145 papers · 4 benchmarks
Set12 is a collection of 12 grayscale images of different scenes that are widely used for evaluation of image denoising methods.
145 papers · 5 benchmarks
VQA-RAD (Visual Question Answering in Radiology)
VQA-RAD consists of 3,515 question–answer pairs on 315 radiology images.
145 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.