Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 247 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 11809–11856 of 12,172
The thickness and appearance of retinal layers are essential markers for diagnosing and studying eye diseases.
0 papers · 0 benchmarks
OSDD (Object State Detection Dataset)
The Objects States Detection Dataset consists of images depicting everyday household objects in a number of different states.
0 papers · 0 benchmarks
OCTCBVS is a benchmark dataset for testing and evaluating novel and state-of-the-art computer vision algorithms.
0 papers · 0 benchmarks
One of the most important aspects of robot scene understanding is semantic segmentation of external environments.
0 papers · 0 benchmarks
OffComBR (Offensive Comments in the Brazilian Web)
Offensive comments obtained from Brazilian website.
0 papers · 0 benchmarks
An annotated data set consisting of user comments posted to an Austrian newspaper website (in German language).
0 papers · 0 benchmarks
Of the 12,330 sessions in the dataset, 84.5% (10,422) were negative class samples that did not end with shopping, and the rest (1908) were positive class samples ending with shopping.
0 papers · 0 benchmarks
The OpenEQA dataset is a significant contribution in the field of Embodied Question Answering (EQA).
0 papers · 0 benchmarks
OpenSurfaces is a large database of annotated surfaces created from real-world consumer photographs.
0 papers · 0 benchmarks
This dataset contains sentences extracted from user reviews on a given topic.
0 papers · 0 benchmarks
This study’s sample consists of seven corporations (Black Rock, Google, Meta, JP Morgan, Walgreens, Netflix, and Pepsico) analyzed across seven quarters beginning in 2021.
0 papers · 0 benchmarks
This dataset is an extremely challenging set of over 2000+ original Oximeter images captured and crowdsourced from over 300+ urban and rural areas, where each image is manually reviewed and verified by computer vision professionals at…
0 papers · 0 benchmarks
Introduce three new neuromorphic vision datasets recorded by a novel neuromorphic vision sensor named Dynamic Vision Sensors (DVS).
0 papers · 0 benchmarks
Data in this study come from western Ecuador's Choco tropical forest, including \textit{Fundación para la Conservación de los Andes Tropicales Reserve and adjacent Reserva Ecológica Mache-Chindul park} (FCAT; 00°23'28'' N, 79°41'05'' W),…
0 papers · 0 benchmarks
PANACEA (PANACEA dataset - Heterogeneous COVID-19 Claims)
The peer-reviewed publication for this dataset has been presented in the 2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), and can be accessed here:…
0 papers · 0 benchmarks
PAVIS RGB-D is a dataset for person re-identification using depth information.
0 papers · 0 benchmarks
PAXRay (PAXRay: A Projected dataset for the segmentation of Anatomical structures in X-Ray data)
Projection of RibFrac CT dataset to a 2D plane to imitate X-Ray data for a total of 880 images with multi-label segmentation masks.
0 papers · 0 benchmarks
The Polish Cyberbullying Dataset is a valuable resource for studying harmful online phenomena, specifically cyberbullying and hate speech in the Polish language.
0 papers · 0 benchmarks
PCN (Pedestrian Color Naming)
Pedestrian Color Naming (PCN) is a dataset for pedestrian color naming, which contains 14,213 images, each of which hand-labeled with color label for each pixel.
0 papers · 0 benchmarks
PCVC (Persian Consonant Vowel Combination)
The Persian Consonant Vowel Combination (PCVC) dataset is a phoneme based speech dataset, and also the first free Persian speech dataset to help Persian speech researchers.
0 papers · 0 benchmarks
This repository is a collection of PDDL generators, some of which have been used to generate benchmarks for the International Planning Competition (IPC).
0 papers · 0 benchmarks
The PEARL dataset comprises with 30K pedestrian images, each annotated with 25 attribute categories, spanning over 146 sub-attributes.
0 papers · 0 benchmarks
Dataset contains annotated photographs of pear fruitlets for object detection tasks using YOLO architecture.
0 papers · 0 benchmarks
We propose a new light field image database called “PINet” inheriting the hierarchical structure from WordNet.
0 papers · 0 benchmarks
The PIROPO database (People in Indoor ROoms with Perspective and Omnidirectional cameras) comprises multiple sequences recorded in two different indoor rooms, using both omnidirectional and perspective cameras.
0 papers · 0 benchmarks
PJM Hourly Energy Consumption Data PJM Interconnection LLC (PJM) is a regional transmission organization (RTO) in the United States.
0 papers · 0 benchmarks
It is a new large-scale pulmonary nodule dataset named PN9, which contains 8,798 thoracic CT scans and a total of 40,439 annotated nodules.
0 papers · 0 benchmarks
POET (Pascal Objects Eye Tracking)
The POET (Pascal Objects Eye Tracking) is a dataset that consists of eye tracking data for the complete trainval set of ten objects classes (cat, dog, bicycle, motorbike, boat, aeroplane, horse, cow, sofa, dining table) from Pascal VOC…
0 papers · 0 benchmarks
POPCORN (POPCORN: Fictional and Synthetic Intelligence Reports for Named Entity Recognition and Relation Extraction Tasks)
POPCORN is a French dataset consisting of 400 validation texts and 400 training texts, all written and annotated manually.
0 papers · 0 benchmarks
PS5k (presentation-slide-pairs)
We introduce a new data set containing 5000 scientific papers and their slides crawled from conference proceeding websites such as aclweb and usenix.
0 papers · 0 benchmarks
Post-Spraying Image Evaluation This dataset is for the paper Deep Learning for Precision Agriculture: Post-Spraying Evaluation and Deposition Estimation (https://arxiv.org/abs/2409.16213).
0 papers · 0 benchmarks
Pan-STARRS (Panoramic Survey Telescope and Rapid Response System (Pan-STARRS))
Pan-STARRS is a system for wide-field astronomical imaging developed and operated by the Institute for Astronomy at the University of Hawaii.
0 papers · 0 benchmarks
The Panoramic Image Database is a panoramic image dataset.
0 papers · 0 benchmarks
This resource contains training and test data for detecting "intent" sentences in email messages.
0 papers · 0 benchmarks
The LAAS Parkour dataset contains 28 RGB videos capturing human subjects performing four typical parkour techniques: safety-vault, kong vault, pull-up and muscle-up.
0 papers · 0 benchmarks
All existing databases of spoofed speech contain attack data that is spoofed in its entirety.
0 papers · 0 benchmarks
This dataset is from Hartmann and Kemmerzell (2010), who, among other things, analyze the causes of the emergence of party ban provisions in sub-Saharan Africa.
0 papers · 0 benchmarks
The Parzival dataset consists of 47 pages by three writers.
0 papers · 0 benchmarks
Pathfinder and Pathfinder-X have proven to be instrumental in training and testing Large Language Models with long-range dependencies.
0 papers · 0 benchmarks
Dataset contains annotated photographs of pear orchard for object detection tasks using YOLO architecture.
0 papers · 0 benchmarks
The data collection took place at 18 intersections in Minnesota.
0 papers · 0 benchmarks
The penguin dataset is a collection of images of penguin colonies in Antarctica coming from the larger penguin watch project, which was setup with the purpose of monitoring their changes in population.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.