Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 65 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 3073–3120 of 12,172
MCMD (Multi-programming-language Commit Message Dataset)
A large-scale dataset in multi-programming languages and with rich information.
10 papers · 0 benchmarks
MCoNaLa is a multilingual dataset to benchmark code generation from natural language commands extending beyond English.
10 papers · 0 benchmarks
MGTAB (Multi-Relational Graph-Based Twitter Account Detection Benchmark)
MGTAB is the first standardized graph-based benchmark for stance and bot detection.
10 papers · 2 benchmarks
The MM-WHS 2017 dataset is a dataset for multi-modality whole heart segmentation.
10 papers · 1 benchmark
MP-DocVQA (Multipage Document Visual Question Answering)
The dataset is aimed to perform Visual Question Answering on multipage industry scanned documents.
10 papers · 0 benchmarks
MUGEN is a large-scale video-audio-text dataset MUGEN, collected using the open-sourced platform game CoinRun.
10 papers · 0 benchmarks
MUStARD++ is a multimodal sarcasm detection dataset (MUStARD) pre-annotated with 9 emotions.
10 papers · 1 benchmark
Memory Maze is a 3D domain of randomized mazes designed for evaluating the long-term memory abilities of RL agents.
10 papers · 0 benchmarks
MeshRIR is a dataset of acoustic room impulse responses (RIRs) at finely meshed grid points.
10 papers · 0 benchmarks
MoGaze is a dataset of full-body motion for everyday manipulation tasks, which includes 1) long sequences of manipulation tasks, 2) the 3D model of the workspace geometry, and 3) eye-gaze.
10 papers · 0 benchmarks
A large-scale dataset that consists of 21,184 claims, where each claim is assigned a truthfulness label and ruling statement, with 58,523 pieces of evidence in the form of text and images.
10 papers · 0 benchmarks
MuSiQue-Ans is a new multihop QA dataset with ~25K 2-4 hop questions using seed questions from 5 existing single-hop datasets.
10 papers · 1 benchmark
Multi-XScience is a large-scale dataset for multi-document summarization of scientific articles.
10 papers · 0 benchmarks
NAF (National Archives Forms Dataset)
This dataset was created with images provided by the United States National Archive and FamilySearch.
10 papers · 0 benchmarks
NLB (Neural Latents Benchmark)
Neural Latents is a benchmark for latent variable modeling of neural population activity.
10 papers · 0 benchmarks
NLI4CT dataset consists of 2,400 annotated statements with accompanying labels, CTRs, and evidence.
10 papers · 0 benchmarks
- NeoRL is a collection of environments and datasets for offline reinforcement learning with a special focus on real-world applications.
10 papers · 0 benchmarks
Nocturne is a 2D, partially observed, driving simulator, built in C++ for speed and exported as a Python library.
10 papers · 0 benchmarks
OCW (Only Connect Wall Dataset and creative problem solving tasks)
The OCW dataset is for evaluating creative problem solving tasks by curating the problems and human performance results from the popular British quiz show Only Connect.
10 papers · 1 benchmark
OLIVES Dataset (Ophthalmic Labels for Investigating Visual Eye Semantics)
Clinical diagnosis of the eye is performed over multifarious data modalities including scalar clinical labels, vectorized biomarkers, two-dimensional fundus images, and three-dimensional Optical Coherence Tomography (OCT) scans.
10 papers · 0 benchmarks
The OU-ISIR Gait Database, Multi-View Large Population Dataset (OU-MVLP) is meant to aid research efforts in the general area of developing, testing and evaluating algorithms for cross-view gait recognition.
10 papers · 1 benchmark
Office-Caltech-10 a standard benchmark for domain adaptation, which consists of Office 10 and Caltech 10 datasets.
10 papers · 1 benchmark
OpenMEVA is a benchmark for evaluating open-ended story generation metrics.
10 papers · 0 benchmarks
The increasing incidence of melanoma has recently promoted the development of computer-aided diagnosis systems for the classification of dermoscopic images.
10 papers · 3 benchmarks
PLAsTiCC (Photometric LSST Astronomical Time-Series Classification Challenge)
The PLAsTiCC dataset is a collection of simulated light curves from the Photometric LSST Astronomical Time-Series Classification Challenge.
10 papers · 0 benchmarks
These are the files containing the Convex Hull and Traveling Salesman Problem dataset present in the “Pointer Networks” paper: Oriol Vinyals, Meire Fortunato, and Navdeep Jaitly.
10 papers · 0 benchmarks
Perception Test is a benchmark designed to evaluate the perception and reasoning skills of multimodal models.
10 papers · 3 benchmarks
PointQA is a set of datasets for Visual Question Datasets (VQA) that require a pointer to an object in the image to be answered correctly.
10 papers · 0 benchmarks
Polyglot-NER builds massive multilingual annotators with minimal human expertise and intervention.
10 papers · 0 benchmarks
QMUL-SurvFace is a surveillance face recognition benchmark that contains 463,507 face images of 15,573 distinct identities captured in real-world uncooperative surveillance scenes over wide space and time.
10 papers · 1 benchmark
REFLACX (Reports and eye-tracking data for localization of abnormalities in chest x-rays)
The REFLACX dataset contains eye-tracking data for 3,032 readings of chest x-rays by five radiologists.
10 papers · 0 benchmarks
REFinD (REFinD: Relation Extraction Financial Dataset)
REFinD is a large-scale annotated dataset of relations, with ∼29K instances and 22 relations amongst 8 types of entity pairs, generated entirely over financial documents.
10 papers · 0 benchmarks
RVSD (Realistic Video DeSnowing Dataset)
Realistic Video DeSnowing Dataset (RVSD) contains a total of 110 pairs of videos.
10 papers · 0 benchmarks
Rainbow is multi-task benchmark for common-sense reasoning that uses different existing QA datasets: aNLI, Cosmos QA, HellaSWAG.
10 papers · 0 benchmarks
Artificial hierarchical datasets to study how neural networks learn hierarchical tasks.
10 papers · 0 benchmarks
ReQA (Retrieval Question-Answering)
Retrieval Question-Answering (ReQA) benchmark tests a model’s ability to retrieve relevant answers efficiently from a large set of documents.
10 papers · 0 benchmarks
ReaSCAN (ReaSCAN: Compositional Reasoning in Language Grounding)
ReaSCAN is a synthetic navigation task that requires models to reason about surroundings over syntactically difficult languages.
10 papers · 0 benchmarks
RealDOF (Single Image Defocus deblurring)
This dataset consists of 50 high resolution image pairs captured by dual-camera setup for single image defocus deblurring .
10 papers · 0 benchmarks
Rico is a public UI corpus with 72K Android UI screens mined from 9.7K Android apps (Deka et al., 2017).
10 papers · 0 benchmarks
RoomR (Room Rearrangement)
The task of Room Rearrangement consists on an agent exploring a room and recording objects' initial configurations.
10 papers · 0 benchmarks
SK-LARGE is a benchmark dataset for object skeleton detection, built on the MS COCO dataset.
10 papers · 1 benchmark
SMD (Server Machine Dataset)
a dataset of time-series anomaly detection
10 papers · 3 benchmarks
SPRSound (SPRSound: Open-Source SJTU Paediatric Respiratory Sound Database)
This repository contains the released respiratory sound database for IEEE BioCAS Respiratory Sound Track Challenges.
10 papers · 0 benchmarks
SST-3 (Stanford Sentiment Treebank: 3-way)
SST-5 is the Stanford Sentiment Treebank 5-way classification dataset (positive, somewhat positive, neutral, somewhat negative, negative).
10 papers · 1 benchmark
SberQuAD (Sberbank Question Answering Dataset)
A large scale analogue of Stanford SQuAD in the Russian language - is a valuable resource that has not been properly presented to the scientific community.
10 papers · 1 benchmark
Significant progress has been made in building generalist robot manipulation policies, yet their scalable and reproducible evaluation remains challenging, as real-world evaluation is operationally expensive and inefficient.
10 papers · 1 benchmark
SkillSpan (Hard and Soft Skill Extraction from English Job Postings)
SkillSpan is a dataset for Skill Extraction (SE).
10 papers · 0 benchmarks
SlowFlow is an optical flow dataset collected by applying Slow Flow technique on data from a high-speed camera and analyzing the performance of the state-of-the-art in optical flow under various levels of motion blur.
10 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.