Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 138 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 6577–6624 of 12,172
Data from the popular Chinese online shopping platform Taobao includes behaviors like buy, add-to-cart, add-to-favorite, and pageview.
2 papers · 1 benchmark
The dataset contains multi-omics data, incuding mRNA, miRNA, and DNA methylation.
2 papers · 1 benchmark
MultiOOD (Multimodal Out-of-Distribution Detection Benchmark)
MultiOOD is the first benchmark for Multimodal OOD Detection and covers diverse dataset sizes and modalities.
2 papers · 0 benchmarks
MultiOpEd is a corpus of multi-perspective news editorials.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
MultiReQA is a cross-domain evaluation for retrieval question answering models.
2 papers · 0 benchmarks
MultiSV is a corpus designed for training and evaluating text-independent multi-channel speaker verification systems.
2 papers · 0 benchmarks
MultiTACRED is a multilingual version of the large-scale TAC Relation Extraction Dataset.
2 papers · 0 benchmarks
Texture-based studies and designs have been in focus recently.
2 papers · 0 benchmarks
The dataset contains training and evaluation data for 12 languages: - Vietnamese - Romanian - Latvian - Czech - Polish - Slovak - Irish - Hungarian - French - Turkish - Spanish - Croatian For each language, one training, one development…
2 papers · 12 benchmarks
The Multinational Structured Address Dataset is a collection of addresses of 61 different countries.
2 papers · 0 benchmarks
Multirotor gym environment for learning control policies for various unmanned aerial vehicles.
2 papers · 0 benchmarks
The original dataset was provided by Orange telecom in France, which contains anonymized and aggregated human mobility data.
2 papers · 0 benchmarks
The MultiviewC dataset mainly contributes to multiview cattle action recognition, 3D objection detection and tracking.
2 papers · 0 benchmarks
The MuseScore dataset is a collection of 344,166 audio and MIDI pairs downloaded from MuseScore website.
2 papers · 0 benchmarks
Music4All-Onion is a large-scale, multi-modal music dataset that expands the Music4All dataset by including 26 additional audio, video, and metadata features for 109,269 music pieces and provides a set of 252,984,396 listening records of…
2 papers · 0 benchmarks
The MusicBrainz20K dataset for entity resolution and entity clustering is based on real records about songs from the MusicBrainz database.
2 papers · 1 benchmark
MyFood Dataset is an image database for segmenting images of Brazilian foods.
2 papers · 0 benchmarks
We applied our framework, dubbed as ”PreNeRF 360”, to enable the use of the Nutrition5k dataset in NeRF and introduce an updated version of this dataset, known as the N5k360 dataset.
2 papers · 0 benchmarks
Experiments on Li-Ion batteries.
2 papers · 1 benchmark
In this competition you will be identifying regions in satellite images that contain certain cloud formations, with label names: Fish, Flower, Gravel, Sugar.
2 papers · 1 benchmark
NBA: This is extended from a Kaggle dataset containing around 400 NBA basketball players.
2 papers · 1 benchmark
The NBA SportVU dataset contains player and ball trajectories for 631 games from the 2015-2016 NBA season.
2 papers · 1 benchmark
NCI (New Corpus for Ireland)
Contains a wide range of texts in Irish, including fiction, news reports, informative texts and official documents.
2 papers · 0 benchmarks
This database offers iris images (with and without contact lenses) of the same eyes captured shortly one after another with illumination coming from two different locations.
2 papers · 0 benchmarks
The archive contains original images from NIH3T3 cells stained with Hoechst 33342 as PNG files.
2 papers · 0 benchmarks
NL-Drive (Nonlinear Autonomous Driving Dataset)
A challenging multi-frame interpolation dataset for autonomous driving scenarios.
2 papers · 1 benchmark
NLI-TR (Natural Language Inference in Turkish)
Natural Language Inference in Turkish (NLI-TR) provides translations of two large English NLI datasets into Turkish and had a team of experts validate their translation quality and fidelity to the original labels.
2 papers · 0 benchmarks
The dataset consists of titles and abstracts from NLP-related papers.
2 papers · 0 benchmarks
This dataset contains charge densities for NMC (Ni, Mn and Co) 2x2x1 supercell (12 transition metal atoms and 12 Li/vacancy site) with varying levels of Li content.
2 papers · 0 benchmarks
NMED-T (Naturalistic Music EEG Dataset - Tempo)
Losorelli, Steven, Nguyen, Duc T., Dmochowski, Jacek P., and Kaneshiro, Blair This dataset contains cortical (EEG) and behavioral data collected during natural music listening.
2 papers · 0 benchmarks
NPSC (Norwegian Parliamentary Speech Corpus)
The Norwegian Parliamentary Speech Corpus (NPSC) is a speech corpus made by the Norwegian Language Bank at the National Library of Norway in 2019-2021.
2 papers · 0 benchmarks
NQiI (Natural Questions In Icelandic)
Natural Questions in Icelandic (NQiI) is a valuable dataset designed for extractive question answering (QA) in the Icelandic language.
2 papers · 0 benchmarks
NQuAD (Nuclear Question Answering Dataset)
NQuAD is a Nuclear Question Answering Dataset, which contains 700+ nuclear Question Answer pairs developed and verified by expert nuclear researchers.
2 papers · 0 benchmarks
A high-quality captured dataset for object relighting.
2 papers · 0 benchmarks
A high-quality synthetic dataset for object relighting.
2 papers · 0 benchmarks
This collection contains images from 422 non-small cell lung cancer (NSCLC) patients.
2 papers · 0 benchmarks
NSD (Natural Scenes Dataset)
The Natural Scenes Dataset (NSD) is a large-scale fMRI dataset conducted at ultra-high-field (7T) strength at the Center of Magnetic Resonance Research (CMRR) at the University of Minnesota.
2 papers · 0 benchmarks
NSVA (NBA dataset for Sports Video Analysis)
NVSA is a large-scale NBA dataset for Sports Video Analysis (NSVA) with a focus on sports video captioning.
2 papers · 0 benchmarks
NTU-X is an extended version of popular NTU dataset.
2 papers · 1 benchmark
The dataset was constructed by first finding suitable publications and then collecting keyphrases from manual annotators.
2 papers · 1 benchmark
NVD (Naturalistic Variation Object Dataset)
Naturalistic Variation Object Dataset (NVD) is a large simulated dataset of 272k images of everyday objects with naturalistic variations such as object pose, scale, viewpoint, lighting and occlusions.
2 papers · 0 benchmarks
NVGaze (NVGaze: An Anatomically-Informed Dataset for Low-Latency, Near-Eye Gaze Estimation)
Quality, diversity, and size of training dataset are critical factors for learning-based gaze estimators.
2 papers · 0 benchmarks
NYT-H is a dataset for distantly-supervised relation extraction, in which DS-labelled training data is used and several annotators to label test data are hired.
2 papers · 0 benchmarks
A collection of over 2,500 novel English words published in the New York Times between November 2017 and March 2019, manually annotated for their class of novelty (such as lexical derivation, dialectal variation, blending, or compounding).
2 papers · 0 benchmarks
The National Lung Screening Trial (NLST) was a randomized controlled trial conducted by the Lung Screening Study group (LSS) and the American College of Radiology Imaging Network (ACRIN) to determine whether screening for lung cancer with…
2 papers · 1 benchmark
The Nations dataset is a small knowledge graph with 14 entities, 55 relations, and 1992 triples describing countries and their political relationships.
2 papers · 0 benchmarks
This dataset consists of 5808 dialogues, based on 2236 unique scenarios.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.