Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 198 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 9457–9504 of 12,172
NGNGAFID-MC consists of over 7500 labeled flights, representing over 11,500 hours of per second flight data recorder readings of 23 sensor parameters.
1 paper · 0 benchmarks
Kinematics Dataset for the NICOL robot (Neuro-inspired Collaborator).
1 paper · 0 benchmarks
The data consists of 21 images of microtubules in PFA-fixed NIH 3T3 mouse embryonic fibroblasts (DSMZ: ACC59) labeled with a mouse anti-alpha-tubulin monoclonal IgG1 antibody (Thermofisher A11126, primary antibody) and visualized by a…
1 paper · 0 benchmarks
NIH-Lymph Node (NIH-LN) contains 388 mediastinal LNs in 90 CT scans and 595 abdominal LNs in 86 scans.
1 paper · 0 benchmarks
NII-CU MAPD (NII-CU Multispectral Aerial Person Detection Dataset)
The National Institute of Informatics - Chiba University (NII-CU) Multispectral Aerial Person Detection Dataset consists of 5,880 pairs of aligned RGB+FIR (Far infrared) images captured from a drone flying at heights between 20 and 50…
1 paper · 2 benchmarks
We announce the release of a new multilingual speaker dataset called NITK-IISc Multilingual Multi-accent Speaker Profiling(NISP) dataset.
1 paper · 0 benchmarks
The NISQA Corpus includes more than 14,000 speech samples with simulated (e.g.
1 paper · 0 benchmarks
NIST Special Database 19 contains NIST's entire corpus of training materials for handprinted document and character recognition.
1 paper · 0 benchmarks
NITEC (Neuro-Information Technology Eye Contact)
This paper introduces NITEC, a new dataset and models for detecting eye contact from an ego-centric camera perspective.
1 paper · 0 benchmarks
NJH is a dataset of over 40,000 tweets about immigration from the US and UK, annotated with six labels for different aspects of incivility and intolerance.
1 paper · 0 benchmarks
A bilingual (English and Chinese natural language queries) dataset which has NL queries annotated with their corresponding GQL queries (i.e.
1 paper · 0 benchmarks
https://arxiv.org/abs/2502.06858
1 paper · 0 benchmarks
The first Portuguese dataset compiled for Native Language Identification (NLI), the task of identifying an author's first language based on their second language writing.
1 paper · 0 benchmarks
NLI4Wills Corpus can be used to train transformers and sentence-transformer models for the validity evaluation of the legal will statements.
1 paper · 0 benchmarks
This project is a collection of three corpora which can be used for evaluating chatbots or other conversational interfaces.
1 paper · 0 benchmarks
NMED-H (Naturalistic Music EEG Dataset - Hindi)
The NMED-H dataset contains scalp EEG responses recorded from 48 adults as they heard intact and scrambled versions of full-length vocal works (Hindi pop songs).
1 paper · 0 benchmarks
NMF (Named Mathematical Formulas)
Mathematical dataset based on 71 famous mathematical identities.
1 paper · 0 benchmarks
NNID (Nearly Nested Image Datasets)
We build what we name the Nearly-Nested Image Datasets (NNID) such that each dataset owns images of the same dimension, and each dataset is issued from a cropped version of the images belonging to the dataset with the biggest dimensions.
1 paper · 0 benchmarks
NOAA/WDS Guyamas Basin (NOAA/WDS Paleoclimatology - Barron et al. 2004 High Resolution Guaymas Basin Geochemical, Diatom, and Silicoflagellate Data)
This archived Paleoclimatology Study is available from the NOAA National Centers for Environmental Information (NCEI), under the World Data Service (WDS) for Paleoclimatology.
1 paper · 0 benchmarks
This dataset contains the full set of experimental waveforms that were used to produce the article "Non-Linear Phase Noise Mitigation over Systems using Constellation Shaping", published in the Journal of Lightwave Technology with DOI:…
1 paper · 0 benchmarks
This corpus contains data files that were generated as part of the NOVIC paper (see above).
1 paper · 0 benchmarks
NPO (Negative and Positive Obstacles)
The dataset is recorded with an on-vehicle ZED stereo camera in both urban and rural environments The dataset contains various lighting conditions, such as normal lights, large-area shadows, dim lights, and sun glare.
1 paper · 1 benchmark
NR-HCPI (Non-redundant Human CPI dataset)
NR-HCPI (Non-redundant Human CPI dataset)
1 paper · 0 benchmarks
To form the collection of nighttime RAW samples, we first selected a total of 150 images with the spatial resolution at 3464×5202 from the training and validation sets provided by the night image challenge.
1 paper · 0 benchmarks
A dataset to encourage research in these environments.
1 paper · 0 benchmarks
Unique radiogenomic dataset from a Non-Small Cell Lung Cancer (NSCLC) cohort of 211 subjects.
1 paper · 0 benchmarks
In the last two years, millions of lives have been lost due to COVID-19.
1 paper · 0 benchmarks
NTPairs (News-Tweet Paired Dataset)
The NTPairs dataset consists of the pairs of news articles and their corresponding tweets that were published by eight media outlets in 2018.
1 paper · 0 benchmarks
Motion similarity annotations for NTU RGB+D 120 dataset to evaluate motion similarity in the real world.
1 paper · 0 benchmarks
NText is an eight million words dataset extracted and preprocessed from nuclear research papers and thesis.
1 paper · 0 benchmarks
Human pose estimation (HPE) in the top-view using fisheye cameras presents a promising and innovative application domain.
1 paper · 0 benchmarks
The NVALT-11 study considered the effect of profylactic brain radiation versus observation in (m=174) patients with advanced non-small cell lung cancer.
1 paper · 0 benchmarks
Te NVALT-8 study (m=200 participants) examined if nadroparin combined with chemotherapy could reduce cancer relapse after surgical removal of a non-small cell lung tumour.
1 paper · 0 benchmarks
This dataset contains 500K high photo-real rendered images of 10 real head models with (yaw, pitch, roll) head pose labels.
1 paper · 0 benchmarks
NYU-VPR is a dataset for Visual place recognition (VPR) that contains more than 200,000 images over a 2km×2km area near the New York University campus, taken within the whole year of 2016.
1 paper · 0 benchmarks
A RGB-D dataset converted from NYUDv2 into COCO-style instance segmentation format.
1 paper · 2 benchmarks
NaSGEC is a new dataset to facilitate research on Chinese grammatical error correction (CGEC) for native speaker texts from multiple domains.
1 paper · 0 benchmarks
Diacritized texts in Modern Hebrew, collected from eleven different sources.
1 paper · 0 benchmarks
A collection of diacritized Hebrew text in a variety of registers and from different sources.
1 paper · 0 benchmarks
Includes co-referent name string pairs along with their similarities.
1 paper · 0 benchmarks
DIT4BEARs Internship Project (at UiT-The Arctic University of Norway) Dataset The dataset contains data of 5 months including weather conditions, friction coefficient, distance traveled, wind speed, surface temperature, air temperature,…
1 paper · 0 benchmarks
The NASA Exoplanet Archive is an online astronomical exoplanet and stellar catalog and data service that collates and cross-correlates astronomical data and information on exoplanets and their host stars, and provides tools to work with…
1 paper · 0 benchmarks
A general purpose text categorization dataset (NatCat) from three online resources: Wikipedia, Reddit, and Stack Exchange.
1 paper · 0 benchmarks
We provide synthetic reflectance, direct shading (shading due to surface geometry and illumination conditions), ambient light and shadow cast ground-truth images.
1 paper · 0 benchmarks
This csv consists of (x-position, y-position, area) tuples of three views (left, middle, right) of downscaled binary masks with aspect ratio kept (64 x 128) from the 2019 YouTube-VIS challenge, which can be found at…
1 paper · 1 benchmark
We scraped the Gutenberg Project and a subset of English Wikipedia to obtain the list of sentences that contain any.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.