Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 141 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 6721–6768 of 12,172
PoKi is a corpus of 61,330 poems written by children from grades 1 to 12.
2 papers · 0 benchmarks
PointPattern is a graph classification dataset constructed by simple point patterns from statistical mechanics.
2 papers · 0 benchmarks
\subsection{Poisson Equation} The Poisson equation with Dirichlet boundary conditions is studied: -Δu = f, in Ω= [0,1]², u = 0, on ∂Ω, where f consists of a Gaussian superposition, with parameters μx,i, μy,i ∼U(0,1) and ∼U(0.025, 0.1).
2 papers · 0 benchmarks
Collects five open polarimetric SAR images, which are images of the San Francisco area.
2 papers · 0 benchmarks
https://huggingface.co/datasets/jdustinwind/Polite
2 papers · 0 benchmarks
The PolyDensity is collected from Polyinfo.
2 papers · 0 benchmarks
Device characteristics data for 835 distinct donor/acceptor systems for polymer solar cells extracted from abstracts of journal papers.
2 papers · 0 benchmarks
Porto Taxi (Taxi Service Trajectory - Prediction Challenge, ECML PKDD 2015)
An accurate dataset describing trajectories performed by all the 442 taxis running in the city of Porto, in Portugal.
2 papers · 0 benchmarks
The PortraitMode-400 dataset is a significant contribution to the field of video recognition, specifically focusing on portrait mode videos.
2 papers · 0 benchmarks
PoseScript is a dataset that pairs a few thousand 3D human poses from AMASS with rich human-annotated descriptions of the body parts and their spatial relationships.
2 papers · 0 benchmarks
The Poser dataset is a dataset for pose estimation which consists of 1927 training and 418 test images.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
Pothole Mix (Pothole Mix Semantic Segmentation Dataset for Road Damage Detection and Segmentation)
This dataset for the semantic segmentation of potholes and cracks on the road surface was assembled from 5 other datasets already publicly available, plus a very small addition of segmented images on our part.
2 papers · 1 benchmark
https://paperswithcode.com/sota/semantic-segmentation-on-isprs-potsdam
2 papers · 2 benchmarks
PragmaticCode is a dataset of real-world open-source Java projects complete with their development environments and dependencies (through their respective build systems).
2 papers · 0 benchmarks
This is the dataset accompanying the paper: "PreBit - A multimodal model with Twitter FinBERT embeddings for extreme price movement prediction of Bitcoin" Zou, Y., & Herremans, D.
2 papers · 0 benchmarks
ProSLU (Profile-based Spoken Language Understanding)
In the paper, to bridge the research gap, we propose a new and important task, Profile-based Spoken Language Understanding (ProSLU), which requires a model not only depends on the text but also on the given supporting profile information.
2 papers · 2 benchmarks
The Prompted Textures Dataset (PTD) is a synthetic texture image dataset consisting of 246,285 images across 56 different texture classes from the work On Synthetic Texture Datasets: Challenges, Creation, and Curation.
2 papers · 0 benchmarks
PropSegmEnt is a corpus of over 35K propositions annotated by expert human raters.
2 papers · 0 benchmarks
A data set introduced for training on the protein design task.
2 papers · 0 benchmarks
ProteinGym is a collection of benchmarks aiming at comparing the ability of models to predict the effects of protein mutations.
2 papers · 0 benchmarks
Psychometric NLP is a corpus for psychometric natural language processing (NLP) related to important dimensions such as trust, anxiety, numeracy, and literacy, in the health domain.
2 papers · 0 benchmarks
PubFig (Public Figures Face Database)
The PubFig database is a large, real-world face dataset consisting of 58,797 images of 200 people collected from the internet.
2 papers · 0 benchmarks
A collection of 385,705 scientific abstracts about Cognitive Control and their GPT-3 embeddings.
2 papers · 1 benchmark
PICO is a framework to formulate a well-defined focused clinical question.
2 papers · 0 benchmarks
PulseImpute is a benchmark for Pulsative Physiological Signal Imputation which includes realistic mHealth missingness models, an extensive set of baselines, and clinically-relevant downstream tasks.
2 papers · 0 benchmarks
The Pump and dump dataset is an annotated set of messages to detect cryptocurrency market manipulations.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
The Pushshift Telegram dataset is made up of over 27.8K channels and 317M messages from 2.2M unique users.
2 papers · 0 benchmarks
Datasets of QCD jets used for studying unfolding in OmniFold: A Method to Simultaneously Unfold All Observables.
2 papers · 0 benchmarks
Python Programming Puzzles (P3) is an open-source dataset where each puzzle is defined by a short Python program , and the goal is to find an input which makes output "True".
2 papers · 0 benchmarks
QDAT data set contains 1500 WAV files along with sound files stored on Excel CSV file format.
2 papers · 0 benchmarks
QMAR (Quality of Movement Assessment for Rehabilitation)
QMAR is an RGB multi-view Quality of Human Movement Assessment dataset.
2 papers · 0 benchmarks
The QTUNA dataset is the result of a series of elicitation experiments in which human speakers were asked to perform a linguistic task that invites the use of quantified expressions in order to inform possible Natural Language Generation…
2 papers · 0 benchmarks
QUITE (Quantifying Uncertainty in natural language Text)
QUITE (Quantifying Uncertainty in natural language Text) is an entirely new benchmark that allows for assessing the capabilities of neural language model-based systems w.r.t.
2 papers · 0 benchmarks
The QV-Pipe dataset consists of 9.6k videos, which are collected from real-world urban pipes.
2 papers · 0 benchmarks
Introduction Generalized quantifiers (e.g., few, most) are used to indicate the proportions predicates are satisfied.
2 papers · 0 benchmarks
Description Quo Vadis: Hybrid Machine Learning Meta-Model Based on Contextual and Behavioral Malware Representations It contains behavioral reports obtained with Speakeasy emulator from 93533 32-bit portable executables (PE).
2 papers · 0 benchmarks
R2VQ (Recipe-to-Video Questions)
R2VQ is a dataset designed for testing competence-based comprehension of machines over a multimodal recipe collection, which contains text-video aligned recipes.
2 papers · 0 benchmarks
RAD (RELEVANCE AND DIVERSITY DATASET)
The dataset is useful for query-adaptive video summarization and annotated with diversity and query-specific relevance labels.
2 papers · 0 benchmarks
RADIOML 2018.01A is a dataset which includes both synthetic simulated channel effects of 24 digital and analog modulation types which has been validated.
2 papers · 0 benchmarks
RAE (Rainforest Automation Energy)
The Rainforest Automation Energy (RAE) dataset was create to help smart grid researchers test their algorithms which make use of smart meter data.
2 papers · 0 benchmarks
A human-revised dataset for seven languages that allows for the evaluation of multilingual RE systems.
2 papers · 0 benchmarks
RETOUCH (RETOUCH -The Retinal OCT Fluid Detection and Segmentation Benchmark and Challenge)
The goal of the challenge is to compare automated algorithms that are able to detect and segment various types of fluids on a common dataset of optical coherence tomography (OCT) volumes representing different retinal diseases, acquired…
2 papers · 0 benchmarks
RFMiD (Retinal Fundus MultiDisease Image Dataset)
According to the WHO, World report on vision 2019, the number of visually impaired people worldwide is estimated to be 2.2 billion, of whom at least 1 billion have a vision impairment that could have been prevented or is yet to be…
2 papers · 0 benchmarks
Used to show systematic performance improvement in applications such as high frame-rate video synthesis, feature/corner detection and tracking, as well as high dynamic range image reconstruction.
2 papers · 0 benchmarks
RGBD1K (A Large-scale Dataset and Benchmark for RGB-D Object Tracking)
RGBD1K is a benchmark for RGB-D Object Tracking which contains 1050 sequences with about 2.5M frames in total.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.