Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 197 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 9409–9456 of 12,172
A detailed description of this dataset can be found in the Zenodo repository: https://zenodo.org/record/7845361#.ZK-k7y9BxhE
1 paper · 0 benchmarks
A detailed description of this dataset can be found in the Zenodo repository: https://zenodo.org/record/8119042#.ZK-jJC9BxhE
1 paper · 0 benchmarks
https://github.com/lingo-iitgn/XME
1 paper · 0 benchmarks
This dataset contains dialogue lines from the games Knights of the Old Republic 1 & 2 and Neverwinter Nights 1.
1 paper · 0 benchmarks
The first annotated corpus for multilingual analysis of potentially unfair clauses in online Terms of Service.
1 paper · 0 benchmarks
Multimedia Goal-oriented Generative Script Learning Dataset This link contains a dataset consisting of multimedia steps for two categories: gardening and crafts.
1 paper · 0 benchmarks
Dataset for multimodal skills assessment focusing on assessing piano player’s skill level.
1 paper · 3 benchmarks
Spectroscopic techniques are essential tools for determining the structure of molecules.
1 paper · 0 benchmarks
Data used in paper entitled "Multiple Testing and Variable Selection along Least Angle Regression's path".
1 paper · 0 benchmarks
Multispectral and HD vineyard orthomosaics from central Portugal - Mulstispectral and HD orthomosaics - 2 distinct vineyards - ground-truth masks for row detection
1 paper · 0 benchmarks
We present a database of multispectral images that were used to emulate the GAP camera.
1 paper · 0 benchmarks
Accompanying expert data and trained models for 2021 IROS paper on Multiview Manipulation.
1 paper · 0 benchmarks
MuscleMap136 is a dataset for video-based Activated Muscle Group Estimation (AMGE) aiming at identifying currently activated muscular regions of humans performing a specific activity.
1 paper · 0 benchmarks
MuseChat Dataset (MuseChat: A Conversational Music Recommendation System for Videos (CVPR 2024 Highlight Paper))
Music recommendation for videos attracts growing interest in multi-modal research.
1 paper · 0 benchmarks
A Crowdsourced Multi-Domain Music Dataset of Europeana Collections The dataset is part of the paper entitled Employing Crowdsourcing for Enriching a Music Knowledge Base in Higher Education presented at the 4th International Conference on…
1 paper · 0 benchmarks
The dataset of images is built upon a collection of 454 samples kindly provided by the urology department of the Hospital Universitary de Bellvitge (Barcelona, Spain) in the time span of several years.
1 paper · 0 benchmarks
This dataset contains information on application install interactions of users in the Myket android application market.
1 paper · 0 benchmarks
M²ConceptBase is a concept-centric multimodal knowledge base designed to bridge the gap between visual and linguistic semantics.
1 paper · 0 benchmarks
N15News is a large-scale multimodal news dataset comprising 200K imagetext pairs and 15 categories, which exceeding the previous news dataset in both the number of categories and samples.
1 paper · 1 benchmark
The training datasets consisting of NACA 4- and 5-digit airfoils, at different flight conditions, were generated using Javafoil.
1 paper · 0 benchmarks
The following datasets: Kidney OCT dataset Three type of tissues were sampled: cortex, medulla, and pelvis.
1 paper · 0 benchmarks
The dataset contains tiles extracted from Whole Slide Images (WSI) of stained tissue samples.
1 paper · 0 benchmarks
NAIST COVID is a multilingual dataset of social media posts related to COVID-19, consisting of microblogs in English and Japanese from Twitter and those in Chinese from Weibo.
1 paper · 0 benchmarks
This dataset contains names that are exclusively associated with a single gender and that have no ambiguous meanings, therefore being exact with respect to both gender and meaning.
1 paper · 0 benchmarks
This dataset extends NAMEXACT by including words that can be used as names, but may not exclusively be used as names in every context.
1 paper · 0 benchmarks
NAO (Natural Adversarial Object)
Natural Adversarial Objects (NAO) is a new dataset to evaluate the robustness of object detection models.
1 paper · 1 benchmark
NAR is a dataset of audio recordings made with the humanoid robot Nao in real world conditions for sound recognition benchmarking.
1 paper · 0 benchmarks
Dataset for our CVPR paper: "ISNAS-DIP: Image-Specific Neural Architecture Search for Deep Image Prior".
1 paper · 0 benchmarks
Extensible Event Stream (XES) software event log obtained through instrumenting the NASA CEV class using the tool available at {https://svn.win.tue.nl/repos/prom/XPort/}.
1 paper · 0 benchmarks
Samples from NASA Perseverance and set of GAN generated synthetic images from Neural Mars.
1 paper · 1 benchmark
The NAVER LABS localization datasets are 5 new indoor datasets for visual localization in challenging real-world environments.
1 paper · 0 benchmarks
NAVVS (Naturalistic audio-visual volumetric sequences)
NAVVS is a volumetric dataset of naturalistic actions whose captured sound and visual appearance yield an open-access resource for immersive and interactive research within an artificial 3D audio-visual environment, such as VR/AR/XR with…
1 paper · 0 benchmarks
The dataset covers the 2022-23 NBA regular season (2022-10-18 to 2023-01-20) which contains 691 games in 92 game days.
1 paper · 0 benchmarks
NBA_Box_Scores_Odds (NBA Team-Level Box Score Statistics (2015-2019), Historical Win Percentages (2014-2018) and Betting Odds (2018/2019))
Dataset Description: NBA Team Statistics, Historical Performance & Betting Odds (2015-2019) Overview This dataset contains team-level box score statistics, historical win percentages, and closing betting odds for NBA games from 2015 to…
1 paper · 0 benchmarks
NBMOD (Noisy Background Multi-Object Dataset for grasp detection)
Introduction NBMOD is a dataset created for researching the task of specific object grasp detection by robots in noisy environments.
1 paper · 1 benchmark
This is a multilabel dataset used for Noise Identification purpose in the paper "A Comparative Analysis of Noise Reduction Methods in Sentiment Analysis on Noisy Bangla Texts" accepted in 2024 The 9th Workshop on Noisy and User-generated…
1 paper · 0 benchmarks
NCANDA (National Consortium on Alcohol and Neurodevelopment in Adolescence)
The NCANDA consortium is composed of an Administrative component at the University of California San Diego, a Data Analysis and Informatics component at SRI International, and five research sites (University of California San Diego, SRI…
1 paper · 0 benchmarks
NCSE v2.0 (NCSE v2.0: A Dataset of OCR-Processed 19th Century English Newspapers)
The NCSE v2.0 is a digitized collection of six 19th-century English periodicals The ground truth contains 358 cropped images of text blocks from 31 pages of 19th century newspaper data
1 paper · 0 benchmarks
NCTE Transcripts consists of 1,660 45-60 minute long 4th and 5th grade elementary mathematics observations collected by the National Center for Teacher Effectiveness (NCTE) between 2010-2013.
1 paper · 0 benchmarks
Business taxonomies automatically constructed from the content of corporate annual reports.
1 paper · 0 benchmarks
NELA-GT-2021 is the fourth installment of the NELA-GT datasets, NELA-GT-2021.
1 paper · 0 benchmarks
This dataset contains spatiotemporal sequences of SST generated by the NEMO ocean engine.
1 paper · 0 benchmarks
NEREL-BIO is an annotation scheme and corpus of PubMed abstracts in Russian and English.
1 paper · 0 benchmarks
NES-VMDB is a dataset containing 98,940 gameplay videos from 389 NES games, each paired with its original soundtrack in symbolic format (MIDI).
1 paper · 0 benchmarks
Data set used in the work One-Shot Recognition of Manufacturing Defects in Steel Surfaces
1 paper · 0 benchmarks
NEWSKVQA is a new dataset of 12K news videos spanning across 156 hours with 1M multiple-choice question-answer pairs covering 8263 unique entities.
1 paper · 0 benchmarks
NEnv (NEnv - Neural Environment Maps)
Dataset of 30 4K HDR environment maps and trained models.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.