12,172 datasets listed, ordered by the archive's paper count. Page 7 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Stanford Online Products (SOP) dataset has 22,634 classes with 120,053 product images.
231 papers · 5 benchmarks
CamVid (Cambridge-driving Labeled Video Database)
CamVid (Cambridge-driving Labeled Video Database) is a road/driving scene understanding database which was originally captured as five video sequences with a 960×720 resolution camera mounted on the dashboard of a car.
227 papers · 4 benchmarks
FlyingThings3D is a synthetic dataset for optical flow, disparity and scene flow estimation.
226 papers · 0 benchmarks
HKU-IS is a visual saliency prediction dataset which contains 4447 challenging images, most of which have either low contrast or multiple salient objects.
226 papers · 3 benchmarks
The Middlebury Stereo dataset consists of high-resolution stereo sequences with complex geometry and pixel-accurate ground-truth disparity data.
223 papers · 5 benchmarks
VisDA-2017 is a simulation-to-real dataset for domain adaptation with over 280,000 images across 12 categories in the training, validation and testing domains.
223 papers · 6 benchmarks
The Vimeo-90K is a large-scale high-quality video dataset for lower-level video processing.
220 papers · 3 benchmarks
ImageNet Long-Tailed is a subset of /dataset/imagenet dataset consisting of 115.8K images from 1000 categories, with maximally 1280 images per class and minimally 5 images per class.
219 papers · 3 benchmarks
DBLP (Citation Network Dataset)
The DBLP is a citation network dataset.
218 papers · 4 benchmarks
DiDeMo (Distinct Describable Moments)
The Distinct Describable Moments (DiDeMo) dataset is one of the largest and most diverse datasets for the temporal localization of events in videos given natural language descriptions.
216 papers · 3 benchmarks
MuST-C currently represents the largest publicly available multilingual corpus (one-to-many) for speech translation.
216 papers · 2 benchmarks
The DUT-OMRON dataset is used for evaluation of Salient Object Detection task and it contains 5,168 high quality images.
214 papers · 4 benchmarks
LAMA (LAnguage Model Analysis)
LAnguage Model Analysis (LAMA) consists of a set of knowledge sources, each comprised of a set of facts.
214 papers · 0 benchmarks
OpenSubtitles is collection of multilingual parallel corpora.
214 papers · 3 benchmarks
Consists of more than 210k videos for 310 audio classes.
211 papers · 3 benchmarks
MegaFace was a publicly available dataset which is used for evaluating the performance of face recognition algorithms with up to a million distractors (i.e., up to a million people who are not in the test set).
210 papers · 3 benchmarks
TrackingNet is a large-scale tracking dataset consisting of videos in the wild.
210 papers · 2 benchmarks
HAM10000 is a dataset of 10000 training images for detecting pigmented skin lesions.
209 papers · 2 benchmarks
FairFace is a face image dataset which is race balanced.
208 papers · 1 benchmark
The data was collected from the English Wikipedia (December 2018).
208 papers · 1 benchmark
AI2 Diagrams (AI2D) is a dataset of over 5000 grade school science diagrams with over 150000 rich annotations, their ground truth syntactic parses, and more than 15000 corresponding multiple choice questions.
207 papers · 1 benchmark
OpenWebText is an open-source recreation of the WebText corpus.
207 papers · 2 benchmarks
The ShanghaiTech Campus dataset has 13 scenes with complex light conditions and camera angles.
207 papers · 4 benchmarks
300W (300 Faces-In-The-Wild)
The 300-W is a face dataset that consists of 300 Indoor and 300 Outdoor in-the-wild images.
206 papers · 9 benchmarks
CIFAR100 few-shots (CIFAR-FS) is randomly sampled from CIFAR-100 (Krizhevsky & Hinton, 2009) by using the same criteria with which miniImageNet has been generated.
206 papers · 2 benchmarks
The NarrativeQA dataset includes a list of documents with Wikipedia summaries, links to full stories, and questions and answers.
206 papers · 1 benchmark
WiC is a benchmark for the evaluation of context-sensitive word embeddings.
206 papers · 3 benchmarks
CodeXGLUE is a benchmark dataset and open challenge for code intelligence.
205 papers · 10 benchmarks
LAION 5B is a large-scale dataset for research purposes consisting of 5,85B CLIP-filtered image-text pairs.
205 papers · 0 benchmarks
MUSAN is a corpus of music, speech and noise.
204 papers · 0 benchmarks
TACRED (The TAC Relation Extraction Dataset)
TACRED is a large-scale relation extraction dataset with 106,264 examples built over newswire and web text from the corpus used in the yearly TAC Knowledge Base Population (TAC KBP) challenges.
204 papers · 2 benchmarks
COCO Captions contains over one and a half million captions describing over 330,000 images.
203 papers · 4 benchmarks
Gowalla is a location-based social networking website where users share their locations by checking-in.
203 papers · 4 benchmarks
Youtube-VOS is a Video Object Segmentation dataset that contains 4,453 videos - 3,471 for training, 474 for validation, and 508 for testing.
203 papers · 10 benchmarks
Clotho is an audio captioning dataset, consisting of 4981 audio samples, and each audio sample has five captions (a total of 24 905 captions).
202 papers · 3 benchmarks
ENZYMES is a dataset of 600 protein tertiary structures obtained from the BRENDA enzyme database.
202 papers · 1 benchmark
The Leeds Sports Pose (LSP) dataset is widely used as the benchmark for human pose estimation.
202 papers · 1 benchmark
VTAB (Visual Task Adaptation Benchmark)
The Visual Task Adaptation Benchmark (VTAB) is a benchmark designed to evaluate general visual representations².
202 papers · 4 benchmarks
Colored MNIST is a synthetic binary classification task derived from MNIST.
201 papers · 0 benchmarks
The HELEN dataset is composed of 2330 face images of 400×400 pixels with labeled facial components generated through manually-annotated contours along eyes, eyebrows, nose, lips and jawline.
201 papers · 1 benchmark
HumanML3D is a 3D human motion-language dataset that originates from a combination of HumanAct12 and Amass dataset.
201 papers · 2 benchmarks
Kvasir-SEG is an open-access dataset of gastrointestinal polyp images and corresponding segmentation masks, manually annotated by a medical doctor and then verified by an experienced gastroenterologist.
201 papers · 2 benchmarks
The IAM database contains 13,353 images of handwritten lines of text created by 657 writers.
198 papers · 1 benchmark
MPI (Max Planck Institute) Sintel is a dataset for optical flow evaluation that has 1064 synthesized stereo images and ground truth data for disparity.
198 papers · 5 benchmarks
PASCAL VOC (PASCAL Visual Object Classes Challenge)
The PASCAL Visual Object Classes (VOC) 2012 dataset contains 20 object categories including vehicles, household, animals, and other: aeroplane, bicycle, boat, bus, car, motorbike, train, bottle, chair, dining table, potted plant, sofa,…
198 papers · 18 benchmarks
YouCook2 is the largest task-oriented, instructional video dataset in the vision community.
198 papers · 7 benchmarks
AISHELL-1 is a corpus for speech recognition research and building speech recognition systems for Mandarin.
197 papers · 1 benchmark
CINIC-10 is a dataset for image classification.
197 papers · 3 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.