12,172 datasets listed, ordered by the archive's paper count. Page 11 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Wizard of Wikipedia is a large dataset with conversations directly grounded with knowledge retrieved from Wikipedia.
145 papers · 1 benchmark
DPED (DSLR Photo Enhancement Dataset)
A large-scale dataset that consists of real photos captured from three different phones and one high-end reflex camera.
144 papers · 0 benchmarks
MedMCQA is a large-scale, Multiple-Choice Question Answering (MCQA) dataset designed to address real-world medical entrance exam questions.
144 papers · 1 benchmark
fMoW (Functional Map of the World)
Functional Map of the World (fMoW) is a dataset that aims to inspire the development of machine learning models capable of predicting the functional purpose of buildings and land use from temporal sequences of satellite images and a rich…
144 papers · 1 benchmark
MPII Human Pose Dataset is a dataset for human pose estimation.
143 papers · 1 benchmark
NABirds V1 is a collection of 48,000 annotated photographs of the 400 species of birds that are commonly observed in North America.
143 papers · 1 benchmark
The ShareGPT4V dataset is a pioneering large-scale resource that features 1.2 million highly descriptive captions.
143 papers · 0 benchmarks
Aff-Wild2 is a large-scale in-the-wild database and an extension of the Aff-Wild dataset for affect recognition.
142 papers · 2 benchmarks
Color BSD68 dataset for image denoising benchmarks is part of The Berkeley Segmentation Dataset and Benchmark.
142 papers · 15 benchmarks
The Flickr30K Entities dataset is an extension to the Flickr30K dataset.
142 papers · 2 benchmarks
The Pix3D dataset is a large-scale benchmark of diverse image-shape pairs with pixel-level 2D-3D alignment.
142 papers · 5 benchmarks
The UCF-Crime dataset is a large-scale dataset of 128 hours of videos.
142 papers · 3 benchmarks
mC4 is a multilingual variant of the C4 dataset called mC4.
142 papers · 0 benchmarks
The One Billion Word dataset is a dataset for language modeling.
141 papers · 2 benchmarks
FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech)
We introduce FLEURS, the Few-shot Learning Evaluation of Universal Representations of Speech benchmark.
141 papers · 1 benchmark
The SummEval dataset is a resource developed by the Yale LILY Lab and Salesforce Research for evaluating text summarization models.
141 papers · 0 benchmarks
CAMO (Camouflaged Object)
Camouflaged Object (CAMO) dataset specifically designed for the task of camouflaged object segmentation.
139 papers · 2 benchmarks
MVBench is a comprehensive Multi-modal Video understanding Benchmark.
139 papers · 3 benchmarks
Multi30K is a large-scale multilingual multimodal dataset for interdisciplinary machine learning research.
139 papers · 2 benchmarks
e-SNLI is used for various goals, such as obtaining full sentence justifications of a model's decisions, improving universal sentence representations and transferring to out-of-domain NLI datasets.
139 papers · 1 benchmark
NSynth is a dataset of one shot instrumental notes, containing 305,979 musical notes with unique pitch, timbre and envelope.
138 papers · 2 benchmarks
The FC100 dataset (Fewshot-CIFAR100) is a newly split dataset based on CIFAR-100 for few-shot learning.
137 papers · 5 benchmarks
NTU RGB+D 120 is a large-scale dataset for RGB+D human action recognition, which is collected from 106 distinct subjects and contains more than 114 thousand video samples and 8 million frames.
137 papers · 8 benchmarks
Oxford5K is the Oxford Buildings Dataset, which contains 5062 images collected from Flickr.
137 papers · 1 benchmark
SEED-Bench consists of 19K multiple choice questions with accurate human annotations (~6 larger than existing benchmarks), which spans 12 evaluation dimensions including the comprehension of both the image and video modality.
137 papers · 0 benchmarks
Cholec80 is an endoscopic video dataset containing 80 videos of cholecystectomy surgeries performed by 13 surgeons.
134 papers · 2 benchmarks
SciERC dataset is a collection of 500 scientific abstract annotated with scientific entities, their relations, and coreference clusters.
134 papers · 7 benchmarks
VIPeR (Viewpoint Invariant Pedestrian Recognition)
The Viewpoint Invariant Pedestrian Recognition (VIPeR) dataset includes 632 people and two outdoor cameras under different viewpoints and light conditions.
134 papers · 0 benchmarks
The “VehicleID” dataset contains CARS captured during the daytime by multiple real-world surveillance cameras distributed in a small city in China.
134 papers · 8 benchmarks
WinoBias contains 3,160 sentences, split equally for development and test, created by researchers familiar with the project.
134 papers · 0 benchmarks
BLUE (Biomedical Language Understanding Evaluation)
The BLUE benchmark consists of five different biomedicine text-mining tasks with ten corpora.
133 papers · 0 benchmarks
The Replay-Attack Database for face spoofing consists of 1300 video clips of photo and video attack attempts to 50 clients, under different lighting conditions.
133 papers · 1 benchmark
RobustBench is a benchmark of adversarial robustness, which as accurately as possible reflects the robustness of the considered models within a reasonable computational budget.
133 papers · 0 benchmarks
SearchQA was built using an in-production, commercial search engine.
133 papers · 1 benchmark
Virtual KITTI is a photo-realistic synthetic video dataset designed to learn and evaluate computer vision models for several video understanding tasks: object detection and multi-object tracking, scene-level and instance-level semantic…
133 papers · 0 benchmarks
ALFWorld contains interactive TextWorld environments (Côté et.
132 papers · 0 benchmarks
CORe50 is a dataset designed for assessing Continual Learning techniques in an Object Recognition context.
132 papers · 0 benchmarks
132 papers · 2 benchmarks
The AlpacaEval set contains 805 instructions form self-instruct, open-assistant, vicuna, koala, hh-rlhf.
131 papers · 2 benchmarks
131 papers · 5 benchmarks
LEVIR-CD is a new large-scale remote sensing building Change Detection dataset.
131 papers · 2 benchmarks
CMU Panoptic is a large scale dataset providing 3D pose annotations (1.5 millions) for multiple people engaging social activities.
131 papers · 4 benchmarks
GoEmotions is a corpus of 58k carefully curated comments extracted from Reddit, with human annotations to 27 emotion categories or Neutral.
130 papers · 0 benchmarks
LIAR is a publicly available dataset for fake news detection.
130 papers · 1 benchmark
LLaVA-Bench is a dataset created to evaluate the capability of large multimodal models (LMM) in more challenging tasks and generalizability to novel domains.
130 papers · 1 benchmark
MSL (Mars Science Laboratory)
This dataset contains expert-labeled telemetry anomaly data from the Mars Science Laboratory (MSL) rover, Curiosity.
130 papers · 1 benchmark
The CityPersons dataset is a subset of Cityscapes which only consists of person annotations.
129 papers · 2 benchmarks
The Make3D dataset is a monocular Depth Estimation dataset that contains 400 single training RGB and depth map pairs, and 134 test samples.
129 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.