Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 203 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 9697–9744 of 12,172
PRECOG (PREdiction of Clinical Outcomes from Genomic Profiles)
The PREdiction of Clinical Outcomes from Genomic profiles (or PRECOG) encompasses 166 cancer expression data sets, including overall survival data for ~18,000 patients diagnosed with 39 distinct malignancies.
1 paper · 0 benchmarks
The Prima head pose dataset consists of 2790 images of 15 persons recorded twice.
1 paper · 1 benchmark
This is the official dataset for PRMBench.
1 paper · 0 benchmarks
PRMC_L2 (Rock, Punk, Metal, and Core - Livehouse Lighting)
Dataset for studying the relationship between music and lighting in live music performances
1 paper · 0 benchmarks
PRONTO (PRONTO heterogeneous benchmark dataset)
The PRONTO heterogeneous benchmark dataset is based on an industrial-scale multiphase flow facility.
1 paper · 1 benchmark
Dataset for automatic pull request title generation.
1 paper · 0 benchmarks
The PS-Eval Dataset is a suite of polysemous and monosemous contexts extracted and filtered from the WiC dataset.
1 paper · 0 benchmarks
A dataset of 18,731 proteins with their PDB code, index of the first residue in their respective DSSP file, their residue sequence and 9-category secondary structure sequence (including polyproline helices).
1 paper · 1 benchmark
PSB2 (The Second Program Synthesis Benchmark Suite)
https://arxiv.org/abs/2106.06086
1 paper · 0 benchmarks
PSM is a financial-domain dataset of the pairwise search matching task.
1 paper · 0 benchmarks
PSU NRTDB (PSU Near-Regular Texture Database)
The PSU Near-Regular Texture Database is a texture dataset.
1 paper · 0 benchmarks
PTCGA200 (Patch TCGA in 200 microns by 512 px)
PTCGA200 is a public pathological H&E image datasets from Patch TCGA in 200 microns by 512 px.
1 paper · 0 benchmarks
PTVD is a plot-oriented multimodal dataset in the TV domain.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
PVDN (Provident Vehicle Detection at Night)
PVDN is a dataset of vehicle detection at night, using light reflections caused by their headlamps.
1 paper · 0 benchmarks
PWISeg (PWISeg Surgical Instruments Dataset)
Overview The Surgical Instruments Recognition Dataset is a groundbreaking collection of high-resolution images (1280x960 pixels) specifically designed for the recognition and categorization of surgical instruments.
1 paper · 0 benchmarks
We introduce a framework for benchmarking multi-step retrosynthesis methods, i.e.
1 paper · 0 benchmarks
PaSa is a dataset to train Machine Learning algorithms to automate the highlighting of patent paragraphs with semantic annotations.
1 paper · 0 benchmarks
This dataset contains one part for the "Subpage-Agnostic Domain Classification" section of our FOCI 2020 paper "Padding Ain’t Enough: Assessing the Privacy Guarantees of Encrypted DNS".
1 paper · 0 benchmarks
This dataset contains the second part of the "Subpage-Agnostic Domain Classification" section of our FOCI 2020 paper "Padding Ain’t Enough: Assessing the Privacy Guarantees of Encrypted DNS".
1 paper · 0 benchmarks
This dataset contains the main data set of our FOCI 2020 paper "Padding Ain’t Enough: Assessing the Privacy Guarantees of Encrypted DNS".
1 paper · 0 benchmarks
PaintNet is a dataset for learning robotic spray painting of free-form 3D objects.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Pan+ChiPhoto dataset is a Chinese character dataset.
1 paper · 0 benchmarks
A dataset of 2D robot recordings with 21 different symbols.
1 paper · 0 benchmarks
Panoramic Video Panoptic Segmentation Dataset is a large-scale dataset that offers high-quality panoptic segmentation labels for autonomous driving.
1 paper · 0 benchmarks
Paper Field is built from the Microsoft Academic Graph and maps paper titles to one of 7 fields of study.
1 paper · 1 benchmark
PapioVoc (Guinea baboon vocalizations dataset automatically extracted with a deep neural network from natural audio recordings)
Abstract The data collection process consisted of continuously recording during one month a group of Guinea baboons living in semi-liberty at the CNRS primatology center in Rousset-sur-Arc (France).
1 paper · 0 benchmarks
Papyrus is comprised of around 60 million data points.
1 paper · 0 benchmarks
We have prepared a dataset, ParagraphOrdreing, which consists of around 300,000 paragraph pairs.
1 paper · 0 benchmarks
The Python dataset introduced in the Parallel Corpus paper (A Parallel Corpus of Python Functions and Documentation Strings for Automated Code Documentation and Code Generation), commonly used for evaluating automated code summarization.
1 paper · 1 benchmark
Anew dataset of facade images from Paris following the Art-deco style.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Parkinson Speech Dataset is an audio dataset consisting of recordings of 20 Parkinson's Disease (PD) patients and 20 healthy subjects.
1 paper · 0 benchmarks
The PD database consists of training and test files.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Despite recent advances in vision-and-language tasks, most progress is still focused on resource-rich languages such as English.
1 paper · 0 benchmarks
The Part-Whole Relations dataset is a dataset of semantic relations between entities.
1 paper · 0 benchmarks
The data set includes information about 120+ elections (configuration settings and descriptive statistics), projects and 125k+ anonymized voters and their budget preferences.
1 paper · 0 benchmarks
PatchDB is a large-scale security patch dataset that contains around 12K security patches and 24K non-security patches from the real world.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
We address the computer-assisted search for prior art by creating a training dataset for supervised machine learning called PatentMatch.
1 paper · 0 benchmarks
PatternCom is a composed image retrieval benchmark based on PatternNet.
1 paper · 1 benchmark
Pavementscapes is a large-scale dataset to develop and evaluate methods for pavement damage segmentation.
1 paper · 0 benchmarks
The PaviaATN data consists of 62 4-channel fluorescence microscopy images of size 2720 × 2720.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Optimization of pedestrian evacuation in different environments
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.