12,172 datasets listed, ordered by the archive's paper count. Page 8 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
WikiQA (Wikipedia open-domain Question Answering)
The WikiQA corpus is a publicly available set of question and sentence pairs, collected and annotated for research on open-domain question answering.
196 papers · 2 benchmarks
Libri-Light is a collection of spoken English audio suitable for training speech recognition systems under limited or no supervision.
194 papers · 2 benchmarks
The Moving MNIST dataset contains 10,000 video sequences, each consisting of 20 frames.
194 papers · 1 benchmark
MNIST-M is created by combining MNIST digits with the patches randomly extracted from color photos of BSDS500 as their background.
193 papers · 1 benchmark
BioASQ (Biomedical Semantic Indexing and Question Answering)
BioASQ is a question answering dataset.
192 papers · 1 benchmark
Daily exchange rates of eight countries’ currencies against the US dollar, spanning from 1990 to 2010 with 7588 timesteps in total.
192 papers · 0 benchmarks
The MOTChallenge datasets are designed for the task of multiple object tracking.
192 papers · 0 benchmarks
BC5CDR (BioCreative V CDR corpus)
BC5CDR corpus consists of 1500 PubMed articles with 4409 annotated chemicals, 5818 diseases and 3116 chemical-disease interactions.
191 papers · 4 benchmarks
The "Flying Chairs" are a synthetic dataset with optical flow ground truth.
191 papers · 0 benchmarks
CMU Multimodal Opinion Sentiment and Emotion Intensity (CMU-MOSEI) is the largest dataset of sentence-level sentiment analysis and emotion recognition in online videos.
190 papers · 3 benchmarks
XQuAD (Cross-lingual Question Answering Dataset) is a benchmark dataset for evaluating cross-lingual question answering performance.
190 papers · 1 benchmark
FewRel (Few-Shot Relation Classification Dataset)
The FewRel (Few-Shot Relation Classification Dataset) contains 100 relations and 70,000 instances from Wikipedia.
189 papers · 3 benchmarks
IHDP (Infant Health and Development Program)
The Infant Health and Development Program (IHDP) is a randomized controlled study designed to evaluate the effect of home visit from specialist doctors on the cognitive test scores of premature infants.
189 papers · 2 benchmarks
LRW (Lip Reading in the Wild)
The Lip Reading in the Wild (LRW) dataset a large-scale audio-visual database that contains 500 different words from over 1,000 speakers.
188 papers · 8 benchmarks
Celeb-DF is a large-scale challenging dataset for deepfake forensics.
187 papers · 0 benchmarks
The CoNLL dataset is a widely used resource in the field of natural language processing (NLP).
187 papers · 35 benchmarks
RESISC45 dataset is a dataset for Remote Sensing Image Scene Classification (RESISC).
187 papers · 3 benchmarks
SGD (Schema-Guided Dialogue)
The Schema-Guided Dialogue (SGD) dataset consists of over 20k annotated multi-domain, task-oriented conversations between a human and a virtual assistant.
186 papers · 2 benchmarks
SUNCG is a large-scale dataset of synthetic 3D scenes with dense volumetric annotations.
186 papers · 0 benchmarks
Argoverse 2 (AV2) is a collection of three datasets for perception and forecasting research in the self-driving domain.
185 papers · 3 benchmarks
The Extended Yale B database contains 2414 frontal-face images with size 192×168 over 38 subjects and about 64 images per subject.
185 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
184 papers · 0 benchmarks
A new benchmark that spans concepts in justice, well-being, duties, virtues, and commonsense morality.
183 papers · 1 benchmark
SA-1B consists of 11M diverse, high resolution, licensed, and privacy protecting images and 1.1B high-quality segmentation masks.
183 papers · 1 benchmark
The SciQ dataset contains 13,679 crowdsourced science exam questions about Physics, Chemistry and Biology, among others.
183 papers · 1 benchmark
OTB-2015, also referred as Visual Tracker Benchmark, is a visual tracking dataset.
182 papers · 1 benchmark
MARS (Motion Analysis and Re-identification Set)
MARS (Motion Analysis and Re-identification Set) is a large scale video based person reidentification dataset, an extension of the Market-1501 dataset.
181 papers · 2 benchmarks
APPS (Automated Programming Progress Standard)
The APPS dataset consists of problems collected from different open-access coding websites such as Codeforces, Kattis, and more.
180 papers · 1 benchmark
HICO-DET is a dataset for detecting human-object interactions (HOI) in images.
180 papers · 4 benchmarks
Dataset of hate speech annotated on Internet forum posts in English at sentence-level.
180 papers · 1 benchmark
IFEval (Instruction Following Evaluation Datset)
This dataset evaluates instruction following ability of large language models.
180 papers · 1 benchmark
Long-range arena (LRA) is an effort toward systematic evaluation of efficient transformer models.
180 papers · 1 benchmark
MORPH is a facial age estimation dataset, which contains 55,134 facial images of 13,617 subjects ranging from 16 to 77 years old.
180 papers · 8 benchmarks
ShapeNetCore is a subset of the full ShapeNet dataset with single clean 3D models and manually verified category and alignment annotations.
180 papers · 1 benchmark
The Breakfast Actions Dataset comprises of 10 actions related to breakfast preparation, performed by 52 different individuals in 18 different kitchens.
179 papers · 6 benchmarks
FUNSD (Form Understanding in Noisy Scanned Documents)
Form Understanding in Noisy Scanned Documents (FUNSD) comprises 199 real, fully annotated, scanned forms.
179 papers · 3 benchmarks
VCR (Visual Commonsense Reasoning)
Visual Commonsense Reasoning (VCR) is a large-scale dataset for cognition-level visual understanding.
179 papers · 13 benchmarks
The WebVision dataset is designed to facilitate the research on learning visual representation from noisy web data.
179 papers · 4 benchmarks
Builds on top of recent data collection efforts by domain experts in these applications and provides a unified collection of datasets with evaluation metrics and train/test splits that are representative of real-world distribution shifts.
179 papers · 0 benchmarks
The AI2’s Reasoning Challenge (ARC) dataset is a multiple-choice question-answering dataset, containing questions from science exams from grade 3 to grade 9.
178 papers · 3 benchmarks
LabelMe database is a large collection of images with ground truth labels for object detection and recognition.
178 papers · 1 benchmark
QuAC (Question Answering in Context)
Question Answering in Context is a large-scale dataset that consists of around 14K crowdsourced Question Answering dialogs with 98K question-answer pairs in total.
178 papers · 1 benchmark
WMT 2016 is a collection of datasets used in shared tasks of the First Conference on Machine Translation.
178 papers · 16 benchmarks
The Hateful Memes data set is a multimodal dataset for hateful meme detection (image + text) that contains 10,000+ new multimodal examples created by Facebook AI.
177 papers · 3 benchmarks
NELL (Never Ending Language Learning)
NELL is a dataset built from the Web via an intelligent agent called Never-Ending Language Learner.
177 papers · 2 benchmarks
PASCAL-5i is a dataset used to evaluate few-shot segmentation.
177 papers · 1 benchmark
Procgen Benchmark includes 16 simple-to-use procedurally-generated environments which provide a direct measure of how quickly a reinforcement learning agent learns generalizable skills.
177 papers · 1 benchmark
The PAMAP2 Physical Activity Monitoring dataset contains data of 18 different physical activities (such as walking, cycling, playing soccer, etc.), performed by 9 subjects wearing 3 inertial measurement units and a heart rate monitor.
176 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.