Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 16 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 721–768 of 12,172
SIM10k is a synthetic dataset containing 10,000 images, which is rendered from the video game Grand Theft Auto V (GTA5).
92 papers · 3 benchmarks
The TGIF-QA dataset contains 165K QA pairs for the animated GIFs from the TGIF dataset [Li et al.
92 papers · 3 benchmarks
The UCSD Anomaly Detection Dataset was acquired with a stationary camera mounted at an elevation, overlooking pedestrian walkways.
92 papers · 4 benchmarks
VITON (VITON-Zalando Dataset)
VITON was a dataset for virtual try-on of clothing items.
92 papers · 1 benchmark
A large-scale multi-object tracking dataset for human tracking in occlusion, frequent crossover, uniform appearance and diverse body gestures.
91 papers · 1 benchmark
FaceWarehouse is a 3D facial expression database that provides the facial geometry of 150 subjects, covering a wide range of ages and ethnic backgrounds.
91 papers · 0 benchmarks
JFLEG (JHU FLuency-Extended GUG corpus)
JFLEG is for developing and evaluating grammatical error correction (GEC).
91 papers · 5 benchmarks
The MIT-States dataset has 245 object classes, 115 attribute classes and ∼53K images.
91 papers · 4 benchmarks
PopQA is an open-domain QA dataset with 14k QA pairs with fine-grained Wikidata entity ID, Wikipedia page views, and relationship type information.
91 papers · 1 benchmark
Spider dataset is used for evaluation in the paper "Structure-Grounded Pretraining for Text-to-SQL".
91 papers · 3 benchmarks
WikiMatrix is a dataset of parallel sentences in the textual content of Wikipedia for all possible language pairs.
91 papers · 0 benchmarks
E2E (End-to-End NLG Challenge)
End-to-End NLG Challenge (E2E) aims to assess whether recent end-to-end NLG systems can generate more complex output by learning from datasets containing higher lexical richness, syntactic complexity and diverse discourse phenomena.
90 papers · 4 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
90 papers · 13 benchmarks
The Oxford-IIIT Pet Dataset has 37 categories with roughly 200 images for each class.
90 papers · 5 benchmarks
ST-VQA (Scene Text Visual Question Answering)
ST-VQA aims to highlight the importance of exploiting high-level semantic information present in images as textual cues in the VQA process.
90 papers · 0 benchmarks
Winoground is a dataset for evaluating the ability of vision and language models to conduct visio-linguistic compositional reasoning.
90 papers · 1 benchmark
The COCO-Text dataset is a dataset for text detection and recognition.
89 papers · 2 benchmarks
The CoNLL-2012 shared task involved predicting coreference in English, Chinese, and Arabic, using the final version, v5.0, of the OntoNotes corpus.
89 papers · 0 benchmarks
ImageNet-O consists of images from classes that are not found in the ImageNet-1k dataset.
89 papers · 0 benchmarks
The MSU-MFSD dataset contains 280 video recordings of genuine and attack faces.
89 papers · 1 benchmark
PathVQA consists of 32,799 open-ended questions from 4,998 pathology images where each question is manually checked to ensure correctness.
89 papers · 0 benchmarks
FaceForensics is a video dataset consisting of more than 500,000 frames containing faces from 1004 videos that can be used to study image or video forgeries.
88 papers · 1 benchmark
Dataset produced for the SAPIEN simulation environment.
88 papers · 0 benchmarks
Logical reasoning is an important ability to examine, analyze, and critically evaluate arguments as they occur in ordinary language as the definition from Law School Admission Council.
88 papers · 4 benchmarks
ASPEC (Asian Scientific Paper Excerpt Corpus)
ASPEC, Asian Scientific Paper Excerpt Corpus, is constructed by the Japan Science and Technology Agency (JST) in collaboration with the National Institute of Information and Communications Technology (NICT).
87 papers · 0 benchmarks
This dataset is for evaluating the performance of intent classification systems in the presence of "out-of-scope" queries, i.e., queries that do not fall into any of the system-supported intent classes.
87 papers · 5 benchmarks
EQA (Embodied Question Answering)
The EQA (Embodied Question Answering) dataset is a dataset of visual questions and answers grounded in House3D.
87 papers · 0 benchmarks
GigaSpeech, an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for semi-supervised and unsupervised…
87 papers · 3 benchmarks
KP20k is a large-scale scholarly articles dataset with 528K articles for training, 20K articles for validation and 20K articles for testing.
87 papers · 3 benchmarks
The first RGB-Thermal urban scene image dataset with pixel-level annotation.
87 papers · 1 benchmark
ONCE (One Million Scenes)
ONCE (One millioN sCenEs) is a dataset for 3D object detection in the autonomous driving scenario.
87 papers · 1 benchmark
PPMI (Parkinson’s Progression Markers Initiative)
The Parkinson’s Progression Markers Initiative (PPMI) dataset originates from an observational clinical and longitudinal study comprising evaluations of people with Parkinson’s disease (PD), those people with high risk, and those who are…
87 papers · 3 benchmarks
UIEB (Underwater Image Enhancement Benchmark Dataset)
Includes 950 real-world underwater images, 890 of which have the corresponding reference images.
87 papers · 1 benchmark
AbstractReasoning is a dataset for abstract reasoning, where the goal is to infer the correct answer from the context panels based on abstract reasoning.
86 papers · 0 benchmarks
DICM is a dataset for low-light enhancement which consists of 69 images collected with commercial digital cameras.
86 papers · 1 benchmark
DIODE (Dense Indoor and Outdoor Depth)
Diode Dense Indoor/Outdoor DEpth (DIODE) is the first standard dataset for monocular depth estimation comprising diverse indoor and outdoor scenes acquired with the same hardware setup.
86 papers · 2 benchmarks
FLoRes-101 is an evaluation benchmark for low-resource and multilingual machine translation.
86 papers · 57 benchmarks
The MovieQA dataset is a dataset for movie question answering.
86 papers · 1 benchmark
SPAQ (Smartphone Photography Attribute and Quality)
The Smartphone Photography Attribute and Quality (SPAQ) dataset is a comprehensive database for the perceptual quality assessment of smartphone photography.
86 papers · 2 benchmarks
VisA (Visual Anomaly Dataset)
The VisA dataset contains 12 subsets corresponding to 12 different objects as shown in the above figure.
86 papers · 3 benchmarks
The Yelp Dataset is a valuable resource for academic research, teaching, and learning.
86 papers · 15 benchmarks
BigEarthNet consists of 590,326 Sentinel-2 image patches, each of which is a section of i) 120x120 pixels for 10m bands; ii) 60x60 pixels for 20m bands; and iii) 20x20 pixels for 60m bands.
85 papers · 3 benchmarks
CAMUS (Cardiac Acquisitions for Multi-structure Ultrasound Segmentation)
This project aims to provide all the materials to the community to resolve the problem of echocardiographic image segmentation and volume estimation from 2D ultrasound sequences (both two and four-chamber views).
85 papers · 0 benchmarks
CULane is a large scale challenging dataset for academic research on traffic lane detection.
85 papers · 1 benchmark
Orkut is a social network dataset consisting of friendship social network and ground-truth communities from Orkut.com on-line social network where users form friendship each other.
85 papers · 0 benchmarks
Structured3D is a large-scale photo-realistic dataset containing 3.5K house designs (a) created by professional designers with a variety of ground truth 3D structure annotations (b) and generate photo-realistic 2D images (c).
85 papers · 7 benchmarks
A large-scale and machine-generated dataset of 274,186 toxic and benign statements about 13 minority groups.
85 papers · 0 benchmarks
CodeContests is a competitive programming dataset for machine-learning.
84 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.