Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 202 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 9649–9696 of 12,172
PAGE (Professional go annotation dataset)
PAGE contains 98,525 games played by 2,007 professional players and spans over 70 years.
1 paper · 0 benchmarks
PAIR-LRT-Human Dataset contains pairs of thermal and RGB images captured using a FLIR Lepton3.5 thermal sensor and a Raspberry Pi camera v2, respectively.
1 paper · 0 benchmarks
The PART-OF dataset is a dataset of relations extracted from a medical ontology.
1 paper · 0 benchmarks
Overview PASSION derm is a pioneering initiative dedicated to closing the diversity gap in dermatology datasets.
1 paper · 0 benchmarks
PC-GITA is a Spanish speech corpus designed to analyze speech impairments in individuals with Parkinson's Disease (PD).
1 paper · 0 benchmarks
PCDS (People Counting Dataset)
Contains over 4,500 videos recorded at the entrance doors of buses in normal and cluttered conditions.
1 paper · 0 benchmarks
In this folder, you will find solutions of the following partial differential equations: - Burgers - Kortweg-de-Vries -Newell-Whitehead - Kuramoto-Sivashinsky You will find more info about how these were generated in the supplementary…
1 paper · 0 benchmarks
PDFM Embeddings are condensed vector representations designed to encapsulate the complex, multidimensional interactions among human behaviors, environmental factors, and local contexts at specific locations.
1 paper · 0 benchmarks
PECC (PECC: Problem Extraction and Coding Challenges)
Recent advancements in large language models (LLMs) have showcased their exceptional abilities across various tasks, such as code generation, problem-solving and reasoning.
1 paper · 1 benchmark
This dataset are about Nafion 112 membrane standard tests and MEA activation tests of PEM fuel cell in various operation condition.
1 paper · 0 benchmarks
We compiled a new dataset (the PERO layout dataset) that contains 683 images from various sources and historical periods with complete manual text block, text line polygon and baseline annotations.
1 paper · 0 benchmarks
PETA-Protein (PETA: Evaluating the Impact of Protein Transfer Learning with Sub-word Tokenization on Downstream Applications)
PETA: Evaluating the Impact of Protein Transfer Learning with Sub-word Tokenization on Downstream Applications
1 paper · 0 benchmarks
PEnG (Pose-Enhanced Geo-Localisation)
This dataset builds upon the SpaGBOL dataset - a graph-based dataset covering numerous cities across the globe for the purpose of structured city-scale Cross-View Geo-Localisation (CVGL).
1 paper · 0 benchmarks
PFN-VT (PFN Visuo-Tactile Dataset)
PFN-VT is a dataset for the estimation of tactile properties from vision, such as slipperiness or roughness.
1 paper · 0 benchmarks
PGDataset (Profile Generation Dataset) is a dataset created for the PGTask (Profile Generation Task), where the goal is to extract/generate a profile sentence given a dialogue utterance.
1 paper · 1 benchmark
First PHP Webshell Opcode Incremental Dataset Motivation To improve the robustness of PHP webshell detection by analyzing low-level opcode patterns, circumventing common code obfuscation and evasion techniques.
1 paper · 0 benchmarks
PIAST (PIAST: A Multimodal Piano Dataset with Audio, Symbolic and Text)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This dataset is well-structured for the physics-informed training of Neural operators for irregular domain geometry, which provides the FEM results of solving a darcy problem in a domain geometry shape of a pentagram.
1 paper · 0 benchmarks
This dataset is well-structured for the physics-informed training of Neural operators for varying domain geometry, which provides the FEM results of solving a darcy problem in a domain geometry shape of a polygon.
1 paper · 0 benchmarks
This dataset is well-structured for the physics-informed training of Neural operators for varying domain geometry, which provides the FEM results of solving a 2D plate stress problem in a domain geometry shape of a rectangle with four…
1 paper · 0 benchmarks
This dataset is well-structured for the physics-informed training of Neural operators for irregular domain geometry, which provides the FEM results of solving a 2D plate stress problem in a domain geometry shape of a rectangle with a hole.
1 paper · 0 benchmarks
We assembled a benchmark of electronic component pinouts, PINS100, containing 100 common parts frequently used in circuits found on high-traffic electronic tutorial websites such as the ARDUINO PROJECT HUB and AUTODESK TINKERCAD CIRCUITS.
1 paper · 0 benchmarks
PIZZA is a dataset for parsing pizza and drink orders, whose semantics cannot be captured by flat slots and intents.
1 paper · 0 benchmarks
PInNED (Personalized Instance-based Navigation Embodied Dataset)
In the last years, the research interest in visual navigation towards objects in indoor environments has grown significantly.
1 paper · 0 benchmarks
A random sample from Pubmed Knowledge Graph.
1 paper · 0 benchmarks
The PKU Sketch Re-ID dataset is constructed by National Engineering Laboratory for Video Technology (NELVT), Peking University.
1 paper · 1 benchmark
Warning: this dataset contains data that may be offensive or harmful.
1 paper · 0 benchmarks
PLAD (Point Line and Depth dataset)
PLAD is a dataset where sparse depth is provided by line-based visual SLAM to verify StructMDC.
1 paper · 1 benchmark
Less complex PLC dataset, the states each have a dedicated feature where the positive flank (0->1 value switch) indicates a state start.
1 paper · 0 benchmarks
More complex PLC dataset, the states each have a unique combination of feature values indicating a state start.
1 paper · 0 benchmarks
PLOD-filtered (PLOD: An Abbreviation Detection Dataset for Scientific Documents)
PLOD: An Abbreviation Detection Dataset This is the PLOD (filtered) Dataset published at LREC 2022.
1 paper · 0 benchmarks
PLOD-unfiltered (PLOD: An Abbreviation Detection Dataset for Scientific Documents)
PLOD: An Abbreviation Detection Dataset This is the PLOD (unfiltered) Dataset published at LREC 2022.
1 paper · 0 benchmarks
Retrieval-based Clinical Decision Support (ReCDS) can aid clinical workflow by providing relevant literature and similar patients for a given patient.
1 paper · 0 benchmarks
PMC-SA (PMC Structured Abstracts)
PMC-SA (PMC Structured Abstracts) is a dataset of academic publications, used for the task of structured summarization.
1 paper · 0 benchmarks
PMPC (Persona Match on Persona-Chat)
PMPC (Persona Match on Persona-Chat) is a dataset for Speaker Persona Detection (SPD) which aims to detect speaker personas based on the plain conversational text.
1 paper · 0 benchmarks
Object detection dataset featuring people walking on grass captured aboard a UAV.
1 paper · 0 benchmarks
POIE (Products for OCR and Information Extraction)
Products for OCR and Information Extraction (POIE) dataset derives from camera images of various products in the real world.
1 paper · 0 benchmarks
The POLARIS dataset is built from a decade of polarimetric observations (2014–2024) conducted with the SPHERE instrument on the Very Large Telescope (VLT).
1 paper · 0 benchmarks
The LiT.RL POLIT-FALSE-n-LEGIT NEWS DB 2016-2017 contains a total of 274 news articles about U.S.
1 paper · 0 benchmarks
A dataset that represents the online media landscape as perceived by an average US news consumer.
1 paper · 0 benchmarks
The POTUS Corpus is a Database of Weekly Addresses for the Study of Stance in Politics and Virtual Agents.
1 paper · 0 benchmarks
PPED (Periodic Phenomena Event-based Dataset)
PPED: Periodic Phenomena Event-based Dataset The dataset features 12 one-second sequences of periodic phenomena (rotation - 01-06, flicker - 07-08, vibration - 09-10 and movement - 11-12) with GT frequencies ranging from 3.2Hz up to 2000Hz…
1 paper · 0 benchmarks
PPFT (Path Planning on Fifty x Thirty)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This dataset contains 23 patients in total.
1 paper · 0 benchmarks
PPTC (PowerPoint Task Completion (PPTC))
Recent evaluations of Large Language Models (LLMs) have centered around testing their zero-shot/few-shot capabilities for basic natural language tasks and their ability to translate instructions into tool APIs.
1 paper · 0 benchmarks
Multitask learning has led to significant advances in Natural Language Processing, including the decaNLP benchmark where question answering is used to frame 10 natural language understanding tasks in a single model.
1 paper · 0 benchmarks
PQAref (Pubmed Question Answering with references)
The PQAref dataset is a dataset for fine-tuning large language models for referenced question-answering in biomedical domain.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.