12,172 datasets listed, ordered by the archive's paper count. Page 13 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
WritingPrompts is a large dataset of 300K human-written stories paired with writing prompts from an online forum.
118 papers · 1 benchmark
AFLW2000-3D is a dataset of 2000 images that have been annotated with image-level 68-point 3D facial landmarks.
117 papers · 8 benchmarks
FLoRes-200 doubles the existing language coverage of FLoRes-101.
117 papers · 1 benchmark
KILT (Knowledge Intensive Language Tasks) is a benchmark consisting of 11 datasets representing 5 types of tasks: Fact-checking (FEVER), Entity linking (AIDA CoNLL-YAGO, WNED-WIKI, WNED-CWEB), Slot filling (T-Rex, Zero Shot RE), Open…
117 papers · 11 benchmarks
The KVASIR Dataset was released as part of the medical multimedia challenge presented by MediaEval.
117 papers · 1 benchmark
Visual Entailment (VE) consists of image-sentence pairs whereby a premise is defined by an image, rather than a natural language sentence as in traditional Textual Entailment tasks.
117 papers · 2 benchmarks
The 20BN-SOMETHING-SOMETHING dataset is a large collection of labeled video clips that show humans performing pre-defined basic actions with everyday objects.
117 papers · 3 benchmarks
A dataset of large scale alignments between Wikipedia abstracts and Wikidata triples.
117 papers · 1 benchmark
VGGFace2 Dataset (Vggface2: A dataset for recognising faces across pose and age)
VGGFace2 is a large-scale face recognition dataset.
117 papers · 0 benchmarks
LLVIP (A Visible-infrared Paired Dataset for Low-light Vision)
Visible-infrared Paired Dataset for Low-light Vision 30976 images (15488 pairs) 24 dark scenes, 2 daytime scenes Support for image-to-image translation (visible to infrared, or infrared to visible), visible and infrared image fusion,…
116 papers · 6 benchmarks
The MRQA (Machine Reading for Question Answering) dataset is a dataset for evaluating the generalization capabilities of reading comprehension systems.
116 papers · 1 benchmark
116 papers · 4 benchmarks
PadChest is a labeled large-scale, high resolution chest x-ray dataset for the automated exploration of medical images along with their associated reports.
116 papers · 0 benchmarks
SciFact is a dataset of 1.4K expert-written claims, paired with evidence-containing abstracts annotated with veracity labels and rationales.
116 papers · 1 benchmark
This corpus comprises of monolingual data for 100+ languages and also includes data for romanized languages.
115 papers · 0 benchmarks
CSIQ (Categorical Subjective Image Quality)
The CSIQ database consists of 30 original images, each is distorted using six different types of distortions at four to five different levels of distortion.
115 papers · 1 benchmark
A labeled benchmark dataset for training machine learning models to statically detect malicious Windows portable executable files.
115 papers · 0 benchmarks
LRS2 (Lip Reading Sentences 2)
The Oxford-BBC Lip Reading Sentences 2 (LRS2) dataset is one of the largest publicly available datasets for lip reading sentences in-the-wild.
115 papers · 10 benchmarks
The Stack contains over 3TB of permissively-licensed source code files covering 30 programming languages crawled from GitHub.
115 papers · 0 benchmarks
highD Dataset (The Highway Drone Dataset Naturalistic Trajectories of 110 500 Vehicles Recorded at German Highways)
The highD dataset is a new dataset of naturalistic vehicle trajectories recorded on German highways.
115 papers · 0 benchmarks
COFW (Caltech Occluded Faces in the Wild)
The Caltech Occluded Faces in the Wild (COFW) dataset is designed to present faces in real-world conditions.
114 papers · 5 benchmarks
MCTest is a freely available set of stories and associated questions intended for research on the machine comprehension of text.
114 papers · 2 benchmarks
QASC (Question Answering via Sentence Composition)
QASC is a question-answering dataset with a focus on sentence composition.
114 papers · 0 benchmarks
Over a period of three years (2009 - 2011) the daily news and weather forecast airings of the German public tv-station PHOENIX featuring sign language interpretation have been recorded and the weather forecasts of a subset of 386 editions…
114 papers · 2 benchmarks
WHAM! (WSJ0 Hipster Ambient Mixtures)
The WSJ0 Hipster Ambient Mixtures (WHAM!) dataset pairs each two-speaker mixture in the wsj0-2mix dataset with a unique noise background scene.
114 papers · 2 benchmarks
AVA (Atomic Visual Actions)
AVA is a project that provides audiovisual annotations of video for improving our understanding of human activity.
113 papers · 7 benchmarks
A repository that contains political events with a specific timestamp.
113 papers · 2 benchmarks
KonIQ-10k (Konstanz Image Quality 10k Database)
KonIQ-10k is a large-scale IQA dataset consisting of 10,073 quality scored images.
113 papers · 2 benchmarks
The NLPR dataset for salient object detection consists of 1,000 image pairs captured by a standard Microsoft Kinect with a resolution of 640×480.
113 papers · 1 benchmark
VOT2016 is a video dataset for visual object tracking.
113 papers · 1 benchmark
VoxPopuli is a large-scale multilingual corpus providing 100K hours of unlabelled speech data in 23 languages.
113 papers · 1 benchmark
WebKB is a dataset that includes web pages from computer science departments of various universities.
113 papers · 2 benchmarks
EgoSchema is very long-form video question-answering dataset, and benchmark to evaluate long video understanding capabilities of modern vision and language systems.
112 papers · 3 benchmarks
GlaS (Gland Segmentation in Colon Histology Images Challenge)
The dataset used in this challenge consists of 165 images derived from 16 H&E stained histological sections of stage T3 or T42 colorectal adenocarcinoma.
112 papers · 1 benchmark
Imagenet32 is a huge dataset made up of small images called the down-sampled version of Imagenet.
112 papers · 4 benchmarks
Visual7W is a large-scale visual question answering (QA) dataset, with object-level groundings and multimodal answers.
112 papers · 1 benchmark
The smallNORB dataset is a datset for 3D object recognition from shape.
112 papers · 1 benchmark
PTC (Predictive Toxicology Challenge)
PTC is a collection of 344 chemical compounds represented as graphs which report the carcinogenicity for rats.
111 papers · 1 benchmark
Reading Comprehension with Commonsense Reasoning Dataset (ReCoRD) is a large-scale reading comprehension dataset which requires commonsense reasoning.
111 papers · 1 benchmark
Wiki-CS is a Wikipedia-based dataset for benchmarking Graph Neural Networks.
111 papers · 1 benchmark
FinQA is a new large-scale dataset with Question-Answering pairs over Financial reports, written by financial experts.
110 papers · 1 benchmark
The M4 dataset is a collection of 100,000 time series used for the fourth edition of the Makridakis forecasting Competition.
110 papers · 0 benchmarks
Dataset Summary Mind2Web is a dataset for developing and evaluating generalist agents for the web that can follow language instructions to complete complex tasks on any website.
110 papers · 1 benchmark
The Neuromorphic-Caltech101 (N-Caltech101) dataset is a spiking version of the original frame-based Caltech101 dataset.
110 papers · 3 benchmarks
OTB2013 is the previous version of the current OTB2015 Visual Tracker Benchmark.
110 papers · 2 benchmarks
PatchCamelyon is an image classification dataset.
110 papers · 4 benchmarks
The Penn Action Dataset contains 2326 video sequences of 15 different actions and human joint annotations for each sequence.
110 papers · 4 benchmarks
The Cross-lingual Choice of Plausible Alternatives (XCOPA) dataset is a benchmark to evaluate the ability of machine learning models to transfer commonsense reasoning across languages.
110 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.