Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 118 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 5617–5664 of 12,172
PRED18 (PRED18: Predator/Prey DAVIS Dataset)
Twenty DAVIS recordings with a total duration of about 1.25 hour were obtained by driving the two robots in the robot arena of the University of Ulster in Londonderry.
3 papers · 0 benchmarks
This dataset contains both infeasible and feasible data points as described in PRIME.
3 papers · 0 benchmarks
PRW-TBPS is a dataset for text based person search task.
3 papers · 0 benchmarks
We collect a total of 13,380 images captured on 2,210 different scenes.
3 papers · 0 benchmarks
The PT Hate Speech is a valuable resource for studying hate speech in the Portuguese language.
3 papers · 0 benchmarks
PTL (Pedestrian-Traffic-Lights)
A dataset of pedestrian traffic lights containing over 5000 photos taken at hundreds of intersections in Shanghai.
3 papers · 0 benchmarks
The Papers with Code Leaderboards dataset is a collection of over 5,000 results capturing performance of machine learning models.
3 papers · 1 benchmark
PWDB (Pulse Wave Database)
Overview This database of simulated arterial pulse waves is designed to be representative of a sample of pulse waves measured from healthy adults.
3 papers · 0 benchmarks
The ability to jointly understand the geometry of objects and plan actions for manipulating them is crucial for intelligent agents.
3 papers · 1 benchmark
This paper presents a benchmark data set for condition monitoring of rolling bearings in combination with an extensive description of the corresponding bearing damage, the data set generation by experiments and results of datadriven…
3 papers · 0 benchmarks
Real-world dataset of ~400 images of cuboid-shaped parcels with full 2D and 3D annotations in the COCO format.
3 papers · 0 benchmarks
Pars-ABSA is a manually annotated Persian dataset, Pars-ABSA, which is verified by 3 native Persian speakers.
3 papers · 0 benchmarks
An open, broad-coverage corpus for informal Persian named entity recognition was collected from Twitter.
3 papers · 0 benchmarks
PatTR (Patent Translation Resource)
PatTR is a sentence-parallel corpus extracted from the MAREC patent collection.
3 papers · 0 benchmarks
Patzig contains handwritten texts written in modern German.
3 papers · 0 benchmarks
Pavia Centre is a hyperspectral dataset acquired by the ROSIS sensor during a flight campaign over Pavia, northern Italy.
3 papers · 0 benchmarks
A corpus of 553k news articles from six Persian news websites and agencies with relatively high quality author extracted keyphrases, which is then filtered and cleaned to achieve higher quality keyphrases.
3 papers · 0 benchmarks
Modeling what makes an advertisement persuasive, i.e., eliciting the desired response from consumer, is critical to the study of propaganda, social psychology, and marketing.
3 papers · 0 benchmarks
PhoNERCOVID19 is a dataset for recognising COVID-19 related named entities in Vietnamese, consisting of 35K entities over 10K sentences.
3 papers · 1 benchmark
The PhotoSynth (PS) dataset for patch matching consists of a total of 30 scenes with 25 scenes for training and 5 scenes for validation.
3 papers · 0 benchmarks
A benchmark for molecular machine learning where improvements in model performance can be immediately observed in the throughput of promising molecules synthesized in the lab.
3 papers · 0 benchmarks
PACS (Physical Audiovisual CommonSense) is the first audiovisual benchmark annotated for physical commonsense attributes.
3 papers · 1 benchmark
This data set consists of over 1500 one- and two-minute EEG recordings, obtained from 109 volunteers [2].
3 papers · 0 benchmarks
Placenta is a benchmark dataset for node classification in an underexplored domain: predicting microanatomical tissue structures from cell graphs in placenta histology whole slide images.
3 papers · 1 benchmark
PoPArt (Poses of People in Art: A Data Set for Human Pose Estimation in Digital Art History)
Throughout the history of art, the pose—as the holistic abstraction of the human body's expression—has proven to be a constant in numerous studies.
3 papers · 1 benchmark
Polyps in the colon are widely known cancer precursors identified by colonoscopy.
3 papers · 1 benchmark
This prostate MRI segmentation dataset is collected from six different data sources.
3 papers · 0 benchmarks
PubChemQA consists of molecules and their corresponding textual descriptions from PubChem.
3 papers · 1 benchmark
PubMedCite is a domain-specific dataset with about 192K biomedical scientific papers and a large citation graph preserving 917K citation relationships between them.
3 papers · 0 benchmarks
Large multimodal models extend the impressive capabilities of large language models by integrating multimodal understanding abilities.
3 papers · 0 benchmarks
PyBullet is an easy to use Python module for physics simulation, robotics and deep reinforcement learning based on the Bullet Physics SDK.
3 papers · 4 benchmarks
Q-Pain, a dataset for assessing bias in medical QA in the context of pain management, one of the most challenging forms of clinical decision-making.
3 papers · 0 benchmarks
QALD-9-Plus Dataset Description QALD-9-Plus is the dataset for Knowledge Graph Question Answering (KGQA) based on well-known QALD-9.
3 papers · 1 benchmark
QST contains 1,167 video clips that are cut out from 216 time-lapse 4K videos collected from YouTube, which can be used for a variety of tasks, such as (high-resolution) video generation, (high-resolution) video prediction,…
3 papers · 0 benchmarks
We consider the problem of referring camouflaged object detection (Ref-COD), a new task that aims to segment specified camouflaged objects based on a small set of referring images with salient target objects.
3 papers · 0 benchmarks
REFreSD (Rationalized English-French Semantic Divergences)
Consists of English-French sentence-pairs annotated with semantic divergence classes and token-level rationales.
3 papers · 0 benchmarks
RESD (Russian Emotional Speech Dialogs with annotated text)
Russian dataset of emotional speech dialogues.
3 papers · 1 benchmark
RETWEET is a dataset of tweets and overall predominant sentiment of their replies.
3 papers · 2 benchmarks
This paper introduces the RGB Arabic Alphabet Sign Language (AASL) dataset.
3 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 0 benchmarks
RLV (Reinforcement Learning with Videos)
We provide video observations of humans performing two simple tasks in natural environments.
3 papers · 0 benchmarks
ROOR is a reading order prediction (ROP) benchmark which annotates layout reading order as ordering relations.
3 papers · 1 benchmark
A new large-scale retail product dataset for fine-grained image classification.
3 papers · 0 benchmarks
RPLAN - a manually collected large-scale densely annotated dataset of floor plans from real residential buildings.
3 papers · 0 benchmarks
RWCP-SSD-Onomatopoeia is a dataset consisting of 155,568 onomatopoeic words paired with audio samples for environmental sound synthesis.
3 papers · 0 benchmarks
We manually labelled 3359 images from the RWTH-PHOENIX-Weather 2014 Development set.
3 papers · 1 benchmark
The RailEye3D dataset, a collection of train-platform scenarios for applications targeting passenger safety and automation of train dispatching, consists of 10 image sequences captured at 6 railway stations in Austria.
3 papers · 0 benchmarks
The RareDis corpus contains more than 5,000 rare diseases and almost 6,000 clinical manifestations are annotated.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.