Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 246 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 11761–11808 of 12,172
MoralChoice is a survey dataset to evaluate the moral beliefs encoded in LLMs.
0 papers · 0 benchmarks
A dataset of all Moroccan money
0 papers · 0 benchmarks
Dataset Description We conducted a BCI experiment for motor imagery movement (MI movement) of the left and right hands with 52 subjects (19 females, mean age ± SD age = 24.8 ± 3.86 years); Each subject took part in the same experiment, and…
0 papers · 0 benchmarks
Dataset description We recruited 15 healthy subjects aged between 22 and 40 years with a mean age of 27 years (standard deviation 5 years).
0 papers · 0 benchmarks
Dataset from the article Evaluation of EEG oscillatory patterns and cognitive process during simple and compound limb motor imagery [1].
0 papers · 0 benchmarks
Dataset from the article A Fully Automated Trial Selection Method for Optimization of Motor Imagery Based Brain-Computer Interface [1].
0 papers · 0 benchmarks
Data Acquisition EEG and NIRS data was collected in an ordinary bright room.
0 papers · 0 benchmarks
The Mouse Embryo Tracking Database is a dataset for tracking mouse embryos.
0 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
0 papers · 0 benchmarks
This dataset is a hash that uses as key the normalized movie name (for instance, {\sf TheAvengers} and as value an array of all the tropes used in that specific movie, as reported by TVTropes.org users.
0 papers · 0 benchmarks
MuS2 (A Real-World Benchmark for Sentinel-2 Multi-Image Super-Resolution)
A real-world dataset for multi-image super-resolution that matches low-resolution Sentinel-2 images with high-resolution WorldView-2 images.
0 papers · 0 benchmarks
Mudestreda (Mudestreda Multimodal Device State Recognition Dataset)
Mudestreda Multimodal Device State Recognition Dataset obtained from real industrial milling device with Time Series and Image Data for Classification, Regression, Anomaly Detection, Remaining Useful Life (RUL) estimation, Signal Drift…
0 papers · 0 benchmarks
This dataset were acquired with the Airphen (Hyphen, Avignon, France) six-band multi-spectral camera configured using the 450/570/675/710/730/850 nm bands with a 10 nm FWHM.
0 papers · 0 benchmarks
Abstract: We introduce the multi-spectral stereo (MS2) outdoor dataset, including stereo RGB, stereo NIR, stereo thermal, stereo LiDAR data, and GPS/IMU information.
0 papers · 0 benchmarks
we propose the augmented KITTI dataset with fog for both camera and LiDAR sensors with different visibility ranges from 20 to 80 meters to best match realistic fog environment.
0 papers · 0 benchmarks
A collection of multilingual sentiment datasets grouped into 3 classes -- positive, neutral, and negative.
0 papers · 0 benchmarks
A great number of situational comedies (sitcoms) are being regularly made and the task of adding laughter tracks to these is a critical task.
0 papers · 0 benchmarks
Multimodal Large Language Models (MLLMs) have shown significant promise in various applications, leading to broad interest from researchers and practitioners alike.
0 papers · 0 benchmarks
BMI/OpenBMI dataset for MI.
0 papers · 0 benchmarks
NEMO (NEMO: A Database for Emotion Analysis Using Functional Near-Infrared Spectroscopy)
We present a dataset for the analysis of human affective states using functional near-infrared spectroscopy (fNIRS).
0 papers · 0 benchmarks
NERGRIT involves machine learning based NLP Tools and a corpus used for Indonesian Named Entity Recognition, Statement Extraction, and Sentiment Analysis.
0 papers · 1 benchmark
A outdoor dataset for UGNA-VPR
0 papers · 0 benchmarks
NHR-Edit (NoHumansRequired Edit Dataset)
NHR-Edit is a training dataset for instruction-based image editing.
0 papers · 0 benchmarks
NIAN (Needle in a Needlestack)
The Needle in a Needlestack (NIAN) is a new benchmark designed to measure how well Language Learning Models (LLMs) pay attention to the information in their context window¹.
0 papers · 0 benchmarks
NIH Grant Abstracts: ExPORTER is a valuable resource for researchers and data enthusiasts.
0 papers · 0 benchmarks
The NKJP-NER dataset is based on a human-annotated part of the National Corpus of Polish (NKJP).
0 papers · 0 benchmarks
NSC (National Speech Corpus)
The National Speech Corpus (NSC) is a significant initiative led by the Info-communications and Media Development Authority (IMDA) of Singapore.
0 papers · 0 benchmarks
NSIDES (Offsides and Twosides (NSIDES v0.1))
Drug side effects and drug-drug interactions were mined from publicly available data.
0 papers · 0 benchmarks
NSMC (Naver Sentiment Movie Corpus)
This is a movie review dataset in the Korean language.
0 papers · 0 benchmarks
Numbers Station Text to SQL
0 papers · 0 benchmarks
NTLNP (wildlife image dataset)
This is an image dataset for object detection of wildlife in the mixed coniferous broad-leaved forest.
0 papers · 0 benchmarks
NTU VIRAL (NTU VIRAL: A Visual-Inertial-Ranging-Lidar Dataset, From an Aerial Vehicle Viewpoint)
In recent years, autonomous robots have become ubiquitous in research and daily life.
0 papers · 0 benchmarks
A vehicle detection database for vision tasks set in the real world.
0 papers · 0 benchmarks
A pedestrian dataset for Person Re-identification.
0 papers · 0 benchmarks
This dataset is flood data in the city of Parepare, South Sulawesi Province, which contains video data collected from social media Instagram.
0 papers · 0 benchmarks
NeoRL-2 includes new task scenarios that better reflect real-world task properties and includes traditional control methods as the data-collecting method.
0 papers · 0 benchmarks
Nepali News Corpus Raw nepali text scrapped from several online websites.
0 papers · 0 benchmarks
A large-scale clothing dataset named NetiLook to discover netizen-style comments.
0 papers · 0 benchmarks
NeuB1 is a microscopic neuronal image dataset for retinal vessel segmentation, which contains 112 images of size 512 x 152.
0 papers · 0 benchmarks
Neural-Code-Search-Evaluation-Dataset presents an evaluation dataset consisting of natural language query and code snippet pairs, with the hope that future work in this area can use this dataset as a common benchmark.
0 papers · 0 benchmarks
Based on RADDLE and SNIPS , we construct Noise-SF, which includes two different perturbation settings.
0 papers · 0 benchmarks
The Nordjylland News dataset is a collection of news articles from Northern Jutland in Denmark.
0 papers · 0 benchmarks
The Nordjylland News summarization dataset is a collection of news articles from the Nordjylland region in Denmark, along with their corresponding summaries.
0 papers · 0 benchmarks
I have listed the 5 best AI nude generators.
0 papers · 0 benchmarks
OASIS-2 (Open Access Series of Imaging Studies)
This set consists of a longitudinal collection of 150 subjects aged 60 to 96.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.