Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 140 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 6673–6720 of 12,172
PASTRIE (Prepositions Annotated with Supersense Tags in Reddit International English)
Prepositions Annotated with Supersense Tags in Reddit International English (PASTRIE) is a new corpus containing manually annotated preposition supersenses of English data from presumed speakers of four L1s: English, French, German, and…
2 papers · 0 benchmarks
PAX-Ray++ (Projected Anatomy in X-Ray Dataset ++)
The PAX-Ray++ dataset uses pseudo-labeled thorax CTs to enable the segmentation of anatomy in Chest X-Rays.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
PCC (Potsdam Commentary Corpus)
The Potsdam Commentary Corpus (PCC) is a corpus of 220 German newspaper commentaries (2.900 sentences, 44.000 tokens) taken from the online issues of the Märkische Allgemeine Zeitung (MAZ subcorpus) and Tagesspiegel (ProCon subcorpus) and…
2 papers · 0 benchmarks
It is a new proposed dataset for point cloud salient object detection that has 2000 training samples and 872 testing samples.
2 papers · 0 benchmarks
PCam200 (Patch Camelyon in 200 microns by 512 px)
PCam200 is a public pathological H&E image dataset from Patch Camelyon in 200 microns by 512 px made in the same manner from Camelyon2016 challenge dataset.
2 papers · 0 benchmarks
PESMOD (PExels Small Moving Object Detection)
The PESMOD (PExels Small Moving Object Detection) dataset consists of high resolution aerial images in which moving objects are labelled manually.
2 papers · 0 benchmarks
The dataset contains 45 documents containing narrative description of business process and their annotations.
2 papers · 0 benchmarks
PETCI (PETCI: A Parallel English Translation Dataset of Chinese Idioms)
PETCI is a Parallel English Translation dataset of Chinese Idioms, collected from an idiom dictionary and Google and DeepL translation.
2 papers · 0 benchmarks
PFD (Playing for Data: Ground Truth from Computer Games)
Recent progress in computer vision has been driven by high-capacity models trained on large datasets.
2 papers · 0 benchmarks
PHANTOM (Physical Anomalous Trajectory or Motion (PHANTOM))
To evaluate the presented approaches, we created the Physical Anomalous Trajectory or Motion (PHANTOM) dataset consisting of six classes featuring everyday objects or physical setups, and showing nine different kinds of anomalies.
2 papers · 1 benchmark
PHSPD (Polarization Human Shape and Pose Dataset)
PHSPD is a home-grown polarization image dataset of various human shapes and poses.
2 papers · 0 benchmarks
PIPPA (Personal Interaction Pairs between People and AI) is a partially-synthetic dataset.
2 papers · 0 benchmarks
The PKU dataset has almost 4,000 images categorized into five groups (G1-G5) that show different situations.
2 papers · 0 benchmarks
The dataset used in the experiments on the paper "Modeling citation worthiness by using attention‑based bidirectional long short‑term memory networks and interpretable models" There are one million sentences in total, and further splitted…
2 papers · 0 benchmarks
The dataset consists of a total of 20 videos, each of which is 5.5 minutes long in duration.
2 papers · 0 benchmarks
POINTREC is a test collection for point of interest (POI) recommendation, comprising of (i) a set of information needs, (ii) a dataset of POIs, and (iii) graded relevance assessments for information need and POI pairs.
2 papers · 0 benchmarks
PPC (Polish Paraphrase Corpus)
The Polish Paraphrase Corpus (PPC) is a dataset consisting of 7000 manually labeled sentence pairs in Polish.
2 papers · 0 benchmarks
PPG-DaLiA is a publicly available dataset for PPG-based heart rate estimation.
2 papers · 0 benchmarks
Automated leaf segmentation is a challenging area in computer vision.
2 papers · 0 benchmarks
PUMaVOS (Partial and Unusual Masks for Video Object Segmentation)
PUMaVOS is a dataset of challenging and practical use cases inspired by the movie production industry.
2 papers · 0 benchmarks
Paint4Poem consists of 301 high-quality poem-painting pairs collected manually from an influential modern Chinese artist Feng Zikai.
2 papers · 0 benchmarks
Dataset Card for The Cancer Genome Atlas (TCGA) Multimodal Dataset The Cancer Genome Atlas (TCGA) Multimodal Dataset is a comprehensive collection of clinical data, pathology reports, molecular, and slide images for cancer patients.
2 papers · 0 benchmarks
Pansharpening Datasets from WorldView 2, WorldView 3, QuickBird, Gaofen 2 sensors.
2 papers · 4 benchmarks
Pano3D is a new benchmark for depth estimation from spherical panoramas.
2 papers · 0 benchmarks
Paper2Fig100k is a dataset with over 100k images of figures and texts from research papers.
2 papers · 0 benchmarks
Used to investigate common crowdsourced paraphrasing issues and for detecting the quality issues.
2 papers · 0 benchmarks
ParaMAWPS (Paraphrased Math Word Problem Solving Repository)
This repository contains the code, data, and models of the paper titled "Math Word Problem Solving by Generating Linguistic Variants of Problem Statements" published in the Proceedings of the 61st Annual Meeting of the Association for…
2 papers · 1 benchmark
To take advantage of the ever-increasing amount of structural data now available, we also trained Paragraph on a larger dataset.
2 papers · 1 benchmark
The Parallel Meaning Bank (PMB), developed at the University of Groningen and building upon the Groningen Meaning Bank, comprises sentences and texts in raw and tokenised format, syntactic analysis, word senses, thematic roles, reference…
2 papers · 0 benchmarks
Parasitic infections have been recognized as one of the most significant causes of illnesses by WHO.
2 papers · 0 benchmarks
Synthetic dataset of over 13,000 images of damaged and intact parcels with full 2D and 3D annotations in the COCO format.
2 papers · 0 benchmarks
The data includes all movement trajectories extracted from the videos of Parkinson's assessments using Convolutional Pose Machines (CPM) as well as the confidence values from CPM.
2 papers · 0 benchmarks
PcMSP is a dataset annotated from 305 open access scientific articles for material science information extraction that simultaneously contains the synthesis sentences extracted from the experimental paragraphs, as well as the entity…
2 papers · 0 benchmarks
PeerSum is a new MDS dataset using peer reviews of scientific publications.
2 papers · 0 benchmarks
PerCQA is the first Persian dataset for CQA (Community Question Answering).
2 papers · 0 benchmarks
Perseus is a dataset for Cross-Lingual Summarization (CLS) which collects about 94K Chinese scientific documents paired with English summaries.
2 papers · 0 benchmarks
The PATIS is a Persian language dataset for intent detection and slot filling.
2 papers · 2 benchmarks
PersonPath22 is a large-scale multi-person tracking dataset containing 236 videos captured mostly from static-mounted cameras, collected from sources where we were given the rights to redistribute the content and participants have given…
2 papers · 1 benchmark
PetFace is a large-scale animal face re-identification dataset that includes 257,484 unique individuals across 13 families and 319 breeds.
2 papers · 0 benchmarks
Glioblastoma-astrocytoma U373 cells on a polyacrylamide substrate Dr.
2 papers · 2 benchmarks
An annotated dataset of 38,800 phishing and benign websites.
2 papers · 0 benchmarks
A new benchmark dataset of webcam images, Photi-LakeIce, from multiple cameras and two different winters, along with pixel-wise ground truth annotations.
2 papers · 0 benchmarks
Introduction The 2016 PhysioNet/CinC Challenge aims to encourage the development of algorithms to classify heart sound recordings collected from a variety of clinical or nonclinical (such as in-home visits) environments.
2 papers · 0 benchmarks
Early Prediction of Sepsis from Clinical Data: The PhysioNet/Computing in Cardiology Challenge 2019 The goal of this Challenge is the early detection of sepsis using physiological data.
2 papers · 0 benchmarks
145k natural language and PDDL problem pairs from the Blocks World, Gripper, and Floor Tile domains.
2 papers · 0 benchmarks
A set of 221 stereo videos captured by the SOCRATES stereo camera trap in a wildlife park in Bonn, Germany between February and July of 2022.
2 papers · 0 benchmarks
PoC (Points of correspondence)
A dataset containing the documents, source and fusion sentences, and human annotations of points of correspondence between sentences.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.