Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 78 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 3697–3744 of 12,172
Type Inference dataset for TypeScript.
7 papers · 1 benchmark
The datasets introduced in Chapter 6 of my PhD thesis are below.
7 papers · 1 benchmark
The Medical Dataset for Abbreviation Disambiguation for Natural Language Understanding (MeDAL) is a large medical text dataset curated for abbreviation disambiguation, designed for natural language understanding pre-training in the medical…
7 papers · 0 benchmarks
The MedDialog dataset (Chinese) contains conversations (in Chinese) between doctors and patients.
7 papers · 0 benchmarks
MegaAge is a large dataset that consists of 41,941 faces annotated with age posterior distributions.
7 papers · 0 benchmarks
MengeROS is an open-source crowd simulation tool for robot navigation that integrates Menge with ROS.
7 papers · 0 benchmarks
MobilityAids is a dataset for perception of people and their mobility aids.
7 papers · 0 benchmarks
MonoPerfCap is a benchmark dataset for human 3D performance capture from monocular video input consisting of around 40k frames, which covers a variety of different scenarios.
7 papers · 0 benchmarks
We propose the MusicQA dataset to train Music-enabled question-answering models and is used for training and evaluating our MU-LLaMA model.
7 papers · 1 benchmark
In MutualFriends, two agents, A and B, each have a private knowledge base, which contains a list of friends with multiple attributes (e.g., name, school, major, etc.).
7 papers · 0 benchmarks
NASA C-MAPSS (Turbofan Engine Degradation Simulation Data Set)
Engine degradation simulation was carried out using C-MAPSS.
7 papers · 2 benchmarks
NDD20 (Northumberland Dolphin Dataset 2020)
Northumberland Dolphin Dataset 2020 (NDD20) is a challenging image dataset annotated for both coarse and fine-grained instance segmentation and categorisation.
7 papers · 0 benchmarks
NES-MDB (Nintendo Entertainment System Music Database)
The Nintendo Entertainment System Music Database (NES-MDB) is a dataset intended for building automatic music composition systems for the NES audio synthesizer.
7 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
7 papers · 1 benchmark
NExT-QA is a VideoQA benchmark targeting the explanation of video contents.
7 papers · 1 benchmark
NIGHTS (Novel Image Generations with Human-Tested Similarity)
A dataset of human similarity judgments over image pairs that are alike in diverse ways.
7 papers · 0 benchmarks
This dataset contains meteorological observations (temperature) at the land-based weather stations located in the United States, collected from the Online Climate Data Directory of the National Oceanic and Atmospheric Administration (NOAA).
7 papers · 1 benchmark
NTU RGB+D 2D is a curated version of NTU RGB+D often used for skeleton-based action prediction and synthesis.
7 papers · 1 benchmark
New3, a set of 527 instances from AMR 3.0, whose original source was the LORELEI DARPA project – not included in the AMR 2.0 training set – consisting of excerpts from newswires and online forum.
7 papers · 1 benchmark
NumtaDB (Assembled Bengali Handwritten Digits)
To benchmark Bengali digit recognition algorithms, a large publicly available dataset is required which is free from biases originating from geographical location, gender, and age.
7 papers · 0 benchmarks
OBP (Open Bandit Dataset)
Open Bandit Dataset is a public real-world logged bandit feedback data.
7 papers · 0 benchmarks
OMMO is a new benchmark for several outdoor NeRF-based tasks, such as novel view synthesis, surface reconstruction, and multi-modal NeRF.
7 papers · 0 benchmarks
ORBIT is a real-world few-shot dataset and benchmark grounded in a real-world application of teachable object recognizers for people who are blind/low vision.
7 papers · 2 benchmarks
OST300 is an outdoor scene dataset with 300 test images of outdoor scenes, and a training set of 7 categories of images with rich textures.
7 papers · 0 benchmarks
The OccludedPASCAL3D+ is a dataset is designed to evaluate the robustness to occlusion for a number of computer vision tasks, such as object detection, keypoint detection and pose estimation.
7 papers · 0 benchmarks
smac+ offensive complicated scenario with 20 parallel episodic buffer.
7 papers · 1 benchmark
SMAC+ offense distant scenario.
7 papers · 1 benchmark
smac+ offensive near scenario with 20 parallel episodic buffer
7 papers · 1 benchmark
Building footprints are useful for a range of important applications, from population estimation, urban planning and humanitarian response, to environmental and climate science.
7 papers · 0 benchmarks
OpenEDS2020 is a dataset of eye-image sequences captured at a frame rate of 100 Hz under controlled illumination, using a virtual-reality head-mounted display mounted with two synchronized eye-facing cameras.
7 papers · 0 benchmarks
OPOSUM is a dataset for the training and evaluation of Opinion Summarization models which contains Amazon reviews from six product domains: Laptop Bags, Bluetooth Headsets, Boots, Keyboards, Televisions, and Vacuums.
7 papers · 0 benchmarks
The Oxford-Affine dataset is a small dataset containing 8 scenes with sequence of 6 images per scene.
7 papers · 0 benchmarks
The Privacy Annotated HMDB51 (PA-HMDB51) dataset is a video-based dataset for evaluating pirvacy protection in visual action recognition algorithms.
7 papers · 0 benchmarks
Over the past few years, different Computer-Aided Diagnosis (CAD) systems have been proposed to tackle skin lesion analysis.
7 papers · 1 benchmark
PARus (Choice of Plausible Alternatives for Russian language)
Choice of Plausible Alternatives for Russian language (PARus) evaluation provides researchers with a tool for assessing progress in open-domain commonsense causal reasoning.
7 papers · 1 benchmark
PDNC (Project Dialogism Novel Corpus)
A annotated dataset of quotations and within-quotation-mentions in 22 full-length English novels.
7 papers · 0 benchmarks
PET (PET: A new Dataset for Process Extraction from Natural Language Text)
The dataset contains 45 documents containing narrative description of business process and their annotations.
7 papers · 0 benchmarks
PFN-PIC (PFN Picking Instructions for Commodities Dataset)
This dataset is a collection of spoken language instructions for a robotic system to pick and place common objects.
7 papers · 0 benchmarks
PHM2017 is a new dataset consisting of 7,192 English tweets across six diseases and conditions: Alzheimer’s Disease, heart attack (any severity), Parkinson’s disease, cancer (any type), Depression (any severity), and Stroke.
7 papers · 0 benchmarks
PIDray is a large-scale dataset which covers various cases in real-world scenarios for prohibited item detection, especially for deliberately hidden items.
7 papers · 0 benchmarks
The PMData dataset aims to combine the traditional lifelogging with sports activity logging.
7 papers · 0 benchmarks
POT-210 (Planar Object Tracking in the Wild: A Benchmark)
Planar object tracking is an actively studied problem in vision-based robotic applications.
7 papers · 0 benchmarks
PROBA-V (PROBA-V Super-Resolution dataset)
The PROBA-V Super-Resolution dataset is the official dataset of ESA's Kelvins competition for "PROBA-V Super Resolution".
7 papers · 1 benchmark
The PS-Battles dataset is gathered from a large community of image manipulation enthusiasts and provides a basis for media derivation and manipulation detection in the visual domain.
7 papers · 0 benchmarks
Electrocardiography (ECG) is a key diagnostic tool to assess the cardiac condition of a patient.
7 papers · 2 benchmarks
Panoptic nuScenes is a benchmark dataset that extends the popular nuScenes dataset with point-wise groundtruth annotations for semantic segmentation, panoptic segmentation, and panoptic tracking tasks.
7 papers · 0 benchmarks
Adopts two subsets of Freebase (Bollacker et al., 2008) as Knowledge Bases to construct the PathQuestion (PQ) and the PathQuestion-Large (PQL) datasets.
7 papers · 1 benchmark
PhoMT is a high-quality and large-scale Vietnamese-English parallel dataset of 3.02M sentence pairs for machine translation.
7 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.