12,172 datasets listed, ordered by the archive's paper count. Page 2 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
MIMIC-III (The Medical Information Mart for Intensive Care III)
The Medical Information Mart for Intensive Care III (MIMIC-III) dataset is a large, de-identified and publicly-available collection of medical records.
1,041 papers · 8 benchmarks
MS MARCO (Microsoft Machine Reading Comprehension Dataset)
The MS MARCO (Microsoft MAchine Reading Comprehension) is a collection of datasets focused on deep learning in search.
1,036 papers · 7 benchmarks
The English Penn Treebank (PTB) corpus, and in particular the section of the corpus corresponding to the articles of Wall Street Journal (WSJ), is one of the most known and used corpus for the evaluation of models for sequence labelling.
1,006 papers · 10 benchmarks
OGB (Open Graph Benchmark)
The Open Graph Benchmark (OGB) is a collection of realistic, large-scale, and diverse benchmark datasets for machine learning on graphs.
1,000 papers · 17 benchmarks
HellaSwag is a challenge dataset for evaluating commonsense NLI that is specially hard for state-of-the-art models, though its questions are trivial for humans (>95% accuracy).
994 papers · 6 benchmarks
The NYU-Depth V2 data set is comprised of video sequences from a variety of indoor scenes as recorded by both the RGB and Depth cameras from the Microsoft Kinect.
986 papers · 16 benchmarks
C4 (Colossal Clean Crawled Corpus)
C4 is a colossal, cleaned version of Common Crawl's web crawl corpus.
981 papers · 1 benchmark
AG News (AG’s News Corpus) is a subdataset of AG's corpus of news articles constructed by assembling titles and description fields of articles from the 4 largest classes (“World”, “Sports”, “Business”, “Sci/Tech”) of AG’s Corpus.
969 papers · 9 benchmarks
The CelebA-HQ dataset is a high-quality version of CelebA that consists of 30,000 images at 1024×1024 resolution.
954 papers · 13 benchmarks
TriviaQA is a realistic text-based question answering dataset which includes 950K question-answer pairs from 662K documents collected from Wikipedia and the web.
953 papers · 5 benchmarks
HotpotQA is a question answering dataset collected on the English Wikipedia, containing about 113K crowd-sourced questions that are constructed to require the introduction paragraphs of two Wikipedia articles to answer.
933 papers · 3 benchmarks
The Flickr30k dataset contains 31,000 images collected from Flickr, together with 5 reference sentences provided by human annotators.
880 papers · 9 benchmarks
Market-1501 is a large-scale public benchmark dataset for person re-identification.
873 papers · 9 benchmarks
DTD (Describable Textures Dataset)
The Describable Textures Dataset (DTD) contains 5640 texture images in the wild.
870 papers · 8 benchmarks
LSUN (Large-scale Scene UNderstanding Challenge)
The Large-scale Scene Understanding (LSUN) challenge aims to provide a different benchmark for large-scale scene classification and understanding.
867 papers · 10 benchmarks
ConceptNet is a knowledge graph that connects words and phrases of natural language with labeled edges.
852 papers · 1 benchmark
The HMDB51 dataset is a large collection of realistic videos from various sources, including movies and web videos.
839 papers · 10 benchmarks
LFW (Labeled Faces in the Wild)
The LFW dataset contains 13,233 images of faces collected from the web.
820 papers · 12 benchmarks
The ActivityNet dataset contains 200 different types of activities and a total of 849 hours of videos collected from YouTube.
807 papers · 17 benchmarks
The Food-101 dataset consists of 101 food categories with 750 training and 250 test images per category, making a total of 101k images.
805 papers · 14 benchmarks
The Stanford Cars dataset consists of 196 classes of cars with a total of 16,185 images, taken from the rear.
790 papers · 13 benchmarks
MRPC (Microsoft Research Paraphrase Corpus)
Microsoft Research Paraphrase Corpus (MRPC) is a corpus consists of 5,801 sentence pairs collected from newswire articles.
786 papers · 4 benchmarks
The Human3.6M dataset is one of the largest motion capture datasets, which consists of 3.6 million human poses and corresponding images captured by a high-speed motion capture system.
783 papers · 13 benchmarks
PIQA (Physical Interaction: Question Answering)
PIQA is a dataset for commonsense reasoning, and was created to investigate the physical knowledge of existing models in NLP.
772 papers · 2 benchmarks
CoNLL-2003 is a named entity recognition dataset released as a part of CoNLL-2003 shared task: language-independent named entity recognition.
755 papers · 6 benchmarks
DomainNet is a dataset of common objects in six different domain.
753 papers · 7 benchmarks
The GQA dataset is a large-scale visual question answering dataset with real images from the Visual Genome dataset and balanced question-answer pairs.
749 papers · 8 benchmarks
IEMOCAP (The Interactive Emotional Dyadic Motion Capture (IEMOCAP) Database)
Multimodal Emotion Recognition IEMOCAP The IEMOCAP dataset consists of 151 videos of recorded dialogues, with 2 speakers per session for a total of 302 videos across the dataset.
749 papers · 3 benchmarks
Audioset is an audio event dataset, which consists of over 2M human-annotated 10-second video clips.
744 papers · 5 benchmarks
DAVIS (Densely Annotated VIdeo Segmentation)
The Densely Annotation Video Segmentation dataset (DAVIS) is a high quality and high resolution densely annotated video segmentation dataset under two resolutions, 480p and 1080p.
734 papers · 10 benchmarks
BSD (Berkeley Segmentation Dataset)
BSD is a dataset used frequently for image denoising and super-resolution.
718 papers · 48 benchmarks
The dataset contains 400 human action classes, with at least 400 video clips for each action.
712 papers · 0 benchmarks
CoLA (Corpus of Linguistic Acceptability)
The Corpus of Linguistic Acceptability (CoLA) consists of 10657 sentences from 23 linguistics publications, expertly annotated for acceptability (grammaticality) by their original authors.
710 papers · 4 benchmarks
The Caltech101 dataset contains images from 101 object categories (e.g., “helicopter”, “elephant” and “chair” etc.) and a background category that contains the images not from the 101 object categories.
709 papers · 10 benchmarks
WinoGrande is a large-scale dataset of 44k problems, inspired by the original WSC design, but adjusted to improve both the scale and the hardness of the dataset.
703 papers · 7 benchmarks
BoolQ (Boolean Questions)
BoolQ is a question answering dataset for yes/no questions containing 15942 examples.
701 papers · 5 benchmarks
The Reddit dataset is a graph dataset from Reddit posts made in the month of September, 2014.
699 papers · 8 benchmarks
Eurosat is a dataset and deep learning benchmark for land use and land cover classification.
687 papers · 8 benchmarks
VoxCeleb1 is an audio dataset containing over 100,000 utterances for 1,251 celebrities, extracted from videos uploaded to YouTube.
680 papers · 10 benchmarks
SemanticKITTI is a large-scale outdoor-scene dataset for point cloud semantic segmentation.
669 papers · 10 benchmarks
PACS (Photo-Art-Cartoon-Sketch)
PACS is an image dataset for domain generalization.
668 papers · 10 benchmarks
MBPP (Mostly Basic Python Programming)
The benchmark consists of around 1,000 crowd-sourced Python programming problems, designed to be solvable by entry-level programmers, covering programming fundamentals, standard library functionality, and so on.
666 papers · 1 benchmark
CLEVR (Compositional Language and Elementary Visual Reasoning)
CLEVR (Compositional Language and Elementary Visual Reasoning) is a synthetic Visual Question Answering dataset.
657 papers · 3 benchmarks
DIV2K is a popular single-image super-resolution dataset which contains 1,000 images with different scenes and is splitted to 800 for training, 100 for validation and 100 for testing.
654 papers · 3 benchmarks
The Office dataset contains 31 object categories in three domains: Amazon, DSLR and Webcam.
643 papers · 7 benchmarks
The FB15k dataset contains knowledge base relation triples and textual mentions of Freebase entity pairs.
641 papers · 6 benchmarks
MSR-VTT (Microsoft Research Video to Text) is a large-scale dataset for the open domain video captioning, which consists of 10,000 video clips from 20 categories, and each video clip is annotated with 20 English sentences by Amazon…
640 papers · 8 benchmarks
OpenBookQA is a new kind of question-answering dataset modeled after open book exams for assessing human understanding of a subject.
635 papers · 3 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.