12,172 datasets listed, ordered by the archive's paper count. Page 9 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
The UCF-QNRF dataset is a crowd counting dataset and it contains large diversity both in scenes, as well as in background types.
176 papers · 1 benchmark
The 3DMATCH benchmark evaluates how well descriptors (both 2D and 3D) can establish correspondences between RGB-D frames of different views.
175 papers · 3 benchmarks
The nocaps benchmark consists of 166,100 human-generated captions describing 15,100 images from the OpenImages validation and test sets.
175 papers · 13 benchmarks
NExT-QA is a VideoQA benchmark targeting the explanation of video contents.
174 papers · 3 benchmarks
R2R is a dataset for visually-grounded natural language navigation in real buildings.
174 papers · 2 benchmarks
SNAP (Stanford Large Network Dataset Collection)
SNAP is a collection of large network datasets.
174 papers · 0 benchmarks
ALFRED (Action Learning From Realistic Environments and Directives)
ALFRED (Action Learning From Realistic Environments and Directives), is a new benchmark for learning a mapping from natural language instructions and egocentric vision to sequences of actions for household tasks.
173 papers · 0 benchmarks
CAMELYON16 (Cancer Metastases in Lymph Nodes Challenge 2016)
The dataset consists of 400 whole-slide images (WSIs) of lymph node sections stained with hematoxylin and eosin (H&E), collected from two medical centers in the Netherlands.
172 papers · 1 benchmark
PAWS-X contains 23,659 human translated PAWS evaluation pairs and 296,406 machine translated training pairs in six typologically distinct languages: French, Spanish, German, Chinese, Japanese, and Korean.
172 papers · 0 benchmarks
RAF-DB (Real-world Affective Faces)
The Real-world Affective Faces Database (RAF-DB) is a dataset for facial expression.
172 papers · 3 benchmarks
ATOMIC is an atlas of everyday commonsense reasoning, organized through 877k textual descriptions of inferential knowledge.
170 papers · 0 benchmarks
The UCY dataset consist of real pedestrian trajectories with rich multi-human interaction scenarios captured at 2.5 Hz (Δt=0.4s).
170 papers · 1 benchmark
LAION-400M is a dataset with CLIP-filtered 400 million image-text pairs, their CLIP embeddings and kNN indices that allow efficient similarity search.
169 papers · 1 benchmark
FER2013 (Facial Expression Recognition 2013 Dataset)
Fer2013 contains approximately 30,000 facial RGB images of different expressions with size restricted to 48×48, and the main labels of it can be divided into 7 types: 0=Angry, 1=Disgust, 2=Fear, 3=Happy, 4=Sad, 5=Surprise, 6=Neutral.
168 papers · 5 benchmarks
SentEval is a toolkit for evaluating the quality of universal sentence representations.
168 papers · 1 benchmark
MLQA (MultiLingual Question Answering)
MLQA (MultiLingual Question Answering) is a benchmark dataset for evaluating cross-lingual question answering performance.
167 papers · 1 benchmark
RLBench is an ambitious large-scale benchmark and learning environment designed to facilitate research in a number of vision-guided manipulation research areas, including: reinforcement learning, imitation learning, multi-task learning,…
167 papers · 2 benchmarks
COD10K (Camouflaged/Concealed Object Detection)
Sensory ecologists have found that this s background matching camouflage strategy works by deceiving the visual perceptual system of the observer.
166 papers · 2 benchmarks
FDDB (Face Detection Dataset and Benchmark)
The Face Detection Dataset and Benchmark (FDDB) dataset is a collection of labeled faces from Faces in the Wild dataset.
165 papers · 1 benchmark
HELM (Holistic Evaluation of Language Models)
The Holistic Evaluation of Language Models (HELM) is a comprehensive framework developed by Stanford University for evaluating foundation language models.
165 papers · 0 benchmarks
CelebAMask-HQ is a large-scale face image dataset that has 30,000 high-resolution face images selected from the CelebA dataset by following CelebA-HQ.
164 papers · 5 benchmarks
The Sports-1M dataset consists of over a million videos from YouTube.
164 papers · 2 benchmarks
The YCB-Video dataset is a large-scale video dataset for 6D object pose estimation.
164 papers · 5 benchmarks
CORD-19 is a free resource of tens of thousands of scholarly articles about COVID-19, SARS-CoV-2, and related coronaviruses for use by the global research community.
163 papers · 1 benchmark
IJB-B (IARPA Janus Benchmark-B)
The IJB-B dataset is a template-based face dataset that contains 1845 subjects with 11,754 images, 55,025 frames and 7,011 videos where a template consists of a varying number of still images and video frames from different sources.
163 papers · 5 benchmarks
SWAG (Situations With Adversarial Generations)
Given a partial description like "she opened the hood of the car," humans can reason about the situation and anticipate what might come next ("then, she examined the engine").
163 papers · 2 benchmarks
YouTubeVIS is a new dataset tailored for tasks like simultaneous detection, segmentation and tracking of object instances in videos and is collected based on the current largest video object segmentation dataset YouTubeVOS.
163 papers · 2 benchmarks
This paper introduces the pipeline to scale the largest dataset in egocentric vision EPIC-KITCHENS.
162 papers · 6 benchmarks
MultiRC (Multi-Sentence Reading Comprehension)
MultiRC (Multi-Sentence Reading Comprehension) is a dataset of short paragraphs and multi-sentence questions, i.e., questions that can be answered by combining information from multiple sentences of the paragraph.
162 papers · 1 benchmark
CrowdHuman is a large and rich-annotated human detection dataset, which contains 15,000, 4,370 and 5,000 images collected from the Internet for training, validation and testing respectively.
161 papers · 2 benchmarks
Objects365 is a large-scale object detection dataset, Objects365, which has 365 object categories over 600K training images.
161 papers · 2 benchmarks
The SALIency in CONtext (SALICON) dataset contains 10,000 training images, 5,000 validation images and 5,000 test images for saliency prediction.
161 papers · 5 benchmarks
CLUSTER is a node classification tasks generated with Stochastic Block Models, which is widely used to model communities in social networks by modulating the intra- and extra-communities connections, thereby controlling the difficulty of…
159 papers · 1 benchmark
MathQA significantly enhances the AQuA dataset with fully-specified operational programs.
159 papers · 1 benchmark
PAWS (Paraphrase Adversaries from Word Scrambling)
Paraphrase Adversaries from Word Scrambling (PAWS) is a dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of…
159 papers · 0 benchmarks
Verbs in COCO (V-COCO) is a dataset that builds off COCO for human-object interaction detection.
159 papers · 1 benchmark
Visual Dialog (VisDial) dataset contains human annotated questions based on images of MS COCO dataset.
159 papers · 6 benchmarks
WSJ0-2mix is a speech recognition corpus of speech mixtures using utterances from the Wall Street Journal (WSJ0) corpus.
159 papers · 3 benchmarks
ELI5 is a dataset for long-form question answering.
158 papers · 1 benchmark
The Materials Project is a collection of chemical compounds labelled with different attributes.
157 papers · 1 benchmark
At the end of 2017 the Civil Comments platform shut down and chose make their ~2m public comments from their platform available in a lasting open archive so that researchers could understand and improve civility in online conversations for…
156 papers · 1 benchmark
IJB-A (IARPA Janus Benchmark A)
The IARPA Janus Benchmark A (IJB-A) database is developed with the aim to augment more challenges to the face recognition task by collecting facial images with a wide variations in pose, illumination, expression, resolution and occlusion.
156 papers · 2 benchmarks
PartNet is a consistent, large-scale dataset of 3D objects annotated with fine-grained, instance-level, and hierarchical 3D part information.
156 papers · 3 benchmarks
A large-scale natural dataset in English to measure stereotypical biases in four domains: gender, profession, race, and religion.
156 papers · 1 benchmark
Total-Text is a text detection dataset that consists of 1,555 images with a variety of text types including horizontal, multi-oriented, and curved text instances.
156 papers · 2 benchmarks
UNSW-NB15 is a network intrusion dataset.
156 papers · 3 benchmarks
ViZDoom is an AI research platform based on the classical First Person Shooter game Doom.
156 papers · 3 benchmarks
DocRED (Document-Level Relation Extraction Dataset) is a relation extraction dataset constructed from Wikipedia and Wikidata.
155 papers · 4 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.