Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 103 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 4897–4944 of 12,172
MuMu is a new dataset of more than 31k albums classified into 250 genre classes.
4 papers · 0 benchmarks
For each dataset we provide a short description as well as some characterization metrics.
4 papers · 0 benchmarks
We introduce MultiScan, a scalable RGBD dataset construction pipeline leveraging commodity mobile devices to scan indoor scenes with articulated objects and web-based semantic annotation interfaces to efficiently annotate object and part…
4 papers · 1 benchmark
MultiSense is a dataset of 9,504 images annotated with an English verb and its translation in Spanish and German.
4 papers · 0 benchmarks
MultiSpider is a large multilingual text-to-SQL dataset which covers seven languages (English, German, French, Spanish, Japanese, Chinese, and Vietnamese).
4 papers · 0 benchmarks
MultiSubs (MultiSubs: A Large-scale Multimodal and Multilingual Dataset)
MultiSubs is a dataset of multilingual subtitles gathered from the OPUS OpenSubtitles dataset, which in turn was sourced from opensubtitles.org.
4 papers · 5 benchmarks
The twitter emoji dataset obtained from CodaLab comprises of 50 thousand tweets along with the associated emoji label.
4 papers · 0 benchmarks
MuseASTE (MuSe-CarASTE: A comprehensive dataset for aspect sentiment triplet extraction in automotive review videos)
•A new benchmark dataset for Aspect Sentiment Triplet Extraction.
4 papers · 1 benchmark
The Musk dataset describes a set of molecules, and the objective is to detect musks from non-musks.
4 papers · 2 benchmarks
The Musk2 dataset is a set of 102 molecules of which 39 are judged by human experts to be musks and the remaining 63 molecules are judged to be non-musks.
4 papers · 1 benchmark
N-Omniglot is a neuromorphic dataset for few-shot learning.
4 papers · 0 benchmarks
NELA-GT-2020 is an updated version of the NELA-GT-2019 dataset.
4 papers · 0 benchmarks
NHA12D (A New Pavement Crack Dataset)
NHA12D is an annotated pavement crack dataset that contains images with different viewpoints and pavements types.
4 papers · 0 benchmarks
NILoc (Neural Inertial Localizatio)
IMU, WiFi data along with aligned Visual SLAM groundtruth locations from a smartphone carried during natural human motion
4 papers · 0 benchmarks
NIPS4Bplus is a richly annotated birdsong audio dataset, that is comprised of recordings containing bird vocalisations along with their active species tags plus the temporal annotations acquired for them.
4 papers · 0 benchmarks
This dataset is an OSN-transmitted (Online Social Network) version of the NIST dataset (https://www.nist.gov/itl/iad/mig/nimble-challenge-2017-evaluation).
4 papers · 1 benchmark
This dataset is an OSN-transmitted (Online Social Network) version of the NIST dataset (https://www.nist.gov/itl/iad/mig/nimble-challenge-2017-evaluation).
4 papers · 1 benchmark
This dataset is an OSN-transmitted (Online Social Network) version of the NIST dataset (https://www.nist.gov/itl/iad/mig/nimble-challenge-2017-evaluation).
4 papers · 1 benchmark
NPHardEval4V is a dynamic reasoning benchmark designed to evaluate the reasoning capabilities of Multimodal Large Language Models (MLLMs).
4 papers · 0 benchmarks
Bike flow data of New York City with grid 16x8.
4 papers · 1 benchmark
Bike flow data of New York City.
4 papers · 1 benchmark
a dataset from A Hierarchical Framework for Relation Extraction with Reinforcement Learning
4 papers · 1 benchmark
NeuralNews is a dataset for machine-generated news detection.
4 papers · 0 benchmarks
The nordland used in SALAD and BoQ (2760 queries, 27592 reference images, threshold: 1 frames).
4 papers · 1 benchmark
The Nottingham Dataset is a collection of 1200 American and British folk songs.
4 papers · 1 benchmark
The dataset used to pre-train NuNER from the NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated Data Contains AI-extracted entities, their concepts, and descriptions from the given text
4 papers · 0 benchmarks
A set of realistic odd-one-out stimuli gathered "in the wild".
4 papers · 0 benchmarks
OASum is a large-scale open-domain aspect-based summarization dataset which contains more than 3.7 million instances with around 1 million different aspects on 2 million Wikipedia pages.
4 papers · 0 benchmarks
OC (Drowsiness-Detection)
These images were generated using UnityEyes simulator, after including essential eyeball physiology elements and modeling binocular vision dynamics.
4 papers · 0 benchmarks
OCD (Out-of-Context Dataset)
OCD (Out-of-Context Dataset) is a synthetic dataset with fine-grained control over scene context.
4 papers · 0 benchmarks
The OLGA dataset contains artist similarities from AllMusic, together with content features from AcousticBrainz.
4 papers · 0 benchmarks
OLPBENCH is a large Open Link Prediction benchmark, which was derived from the state-of-the-art Open Information Extraction corpus OPIEC (Gashteovski et al., 2019).
4 papers · 0 benchmarks
OMD (Oxford Multimotion Dataset)
The Oxford Multimotion Dataset (OMD) provides a number of multimotion estimation problems of varying complexity.
4 papers · 0 benchmarks
OVDEval includes 9 sub-tasks and introduces evaluations on commonsense knowledge, attribute understanding, position understanding, object relation comprehension, and more.
4 papers · 0 benchmarks
OVQA contains 19,020 medical visual question and answer pairs generated from 2,001 medical images collected from 2,212 EMRs in Orthopedics.
4 papers · 0 benchmarks
Synthetic omnidirectional multi-view image dataset.
4 papers · 0 benchmarks
OpenViVQA (Open-domain Visual Question Answering in Vietnamese)
In recent years, visual question answering (VQA) has attracted attention from the research community because of its highly potential applications (such as virtual assistance on intelligent cars, assistant devices for blind people, or…
4 papers · 0 benchmarks
Oracle-MNIST (Oracle-MNIST: a Realistic Image Dataset for Benchmarking Machine Learning Algorithms)
We introduce the Oracle-MNIST dataset, comprising of 2828 grayscale images of 30,222 ancient characters from 10 categories, for benchmarking pattern classification, with particular challenges on image noise and distortion.
4 papers · 1 benchmark
We introduced this dataset in Points2Surf, a method that turns point clouds into meshes.
4 papers · 0 benchmarks
This is the set of instances use in the PACE 2018 competition, of optimal Steiner Tree computation.
4 papers · 0 benchmarks
PADv2 (Purpose-driven Affordance Dataset v2)
With complex scenes and rich annotations, the PADv2 dataset can be used as a test bed to benchmark affordance detection methods and may also facilitate downstream vision tasks, such as scene understanding, action recognition, and robot…
4 papers · 0 benchmarks
From PARIS: Part-level Reconstruction and Motion Analysis for Articulated Objects: 5.1.
4 papers · 0 benchmarks
PASTEL is a parallelly annotated stylistic language dataset.
4 papers · 0 benchmarks
PASTIS-R (Panoptic Segmentation of Radar and Optical Satellite image TIme Series)
Extension of the PASTIS benchmark with radar and optical image time series.
4 papers · 2 benchmarks
PBC (Mayo Clinic Primary Biliary Cholangitis data)
Primary sclerosing cholangitis is an autoimmune disease leading to destruction of the small bile ducts in the liver.
4 papers · 0 benchmarks
PCBA dataset 11 is a collection of high-quality dose-response data, formulated as a multitask learning benchmark from 128 high-throughput screening (HTS) assays.
4 papers · 2 benchmarks
PCFG SET (Probabilistic Context Free Grammar String Edit Task)
The Probabilistic Context Free Grammar String Edit Task (PCFG SET) dataset is a dataset with sequence to sequence problems specifically designed to test different aspects of compositional generalisation.
4 papers · 0 benchmarks
PDEBench provides a diverse and comprehensive set of benchmarks for scientific machine learning, including challenging and realistic physical problems.
4 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.