Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 205 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 9793–9840 of 12,172
Advanced pixel shift technology is employed to perform a full color sampling of the image.
1 paper · 0 benchmarks
Placepedia contains 240K places with 35M images from all over the world.
1 paper · 0 benchmarks
PlainFact is a high-quality human-annotated dataset with fine-grained explanation (i.e., added information) annotations.
1 paper · 0 benchmarks
The dataset consists of 90 000 color videos that show a planar robot manipulator executing articulated manipulation tasks.
1 paper · 0 benchmarks
An evaluation dataset for planning with LLM agents
1 paper · 0 benchmarks
See the description in the github repository
1 paper · 0 benchmarks
We established a large-scale plant disease segmentation dataset named PlantSeg.
1 paper · 0 benchmarks
Platinum Benchmarks are benchmarks that are are carefully curated to minimize label errors and ambiguity, allowing us to measure reliability of models.
1 paper · 0 benchmarks
A dataset containing four sets of playing card images.
1 paper · 0 benchmarks
PoTATO is a dataset designed to enhance the detection of floating plastic waste in aquatic environments by leveraging polarimetric imaging.
1 paper · 0 benchmarks
The PointDenoisingBenchmark dataset features 28 different shapes, split into 18 training shapes and 10 test shapes.
1 paper · 0 benchmarks
This repository contains a dataset and machine learning algorithms to detect poisoned water from clean water via using equivalent Smartphone embedded Wi-Fi CSI data.
1 paper · 0 benchmarks
Poker Hand Histories A collection of poker hand histories, covering 11 poker variants, in the poker hand history (PHH) format.
1 paper · 0 benchmarks
PolarRR is a new dataset with more than 100 types of glass in which obtained transmission images are perfectly aligned with input mixed images.
1 paper · 0 benchmarks
The dataset includes polarimetric, RGB and depth automotive (on the road) data.
1 paper · 0 benchmarks
The current industrial pipeline includes 315 dynamic industrial scenarios, which can be categorized into three types: QR codes, text, and products.
1 paper · 0 benchmarks
Engagement with the government of Taiwan as part of the vTaiwan participatory process which led to the successful regulation of Uber in Taiwan.
1 paper · 0 benchmarks
Consists of 10K sentence pairs which are human-annotated for semantic relatedness and entailment.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
A dataset for detecting specific text chunks and categories of political advertising in the Polish language.
1 paper · 0 benchmarks
TPM values together with cell type annotations that were obtained from Alex Pollen on 15/10/15 Source: Low-coverage single-cell mRNA sequencing reveals cellular heterogeneity and activated signaling pathways in developing cerebral cortex
1 paper · 1 benchmark
Poly-FEVER is a multilingual fact verification benchmark designed to evaluate hallucination detection in large language models (LLMs).
1 paper · 0 benchmarks
PolyMATH, a challenging benchmark aimed at evaluating the general cognitive reasoning abilities of MLLMs.
1 paper · 0 benchmarks
PolyNews is a multilingual dataset containing news titles in 77 languages and 19 scripts.
1 paper · 0 benchmarks
PolyNews is a multilingual parallel dataset containing news titles 833 language pairs, spanning in 64 languages and 17 scripts.
1 paper · 0 benchmarks
PolyU-BPCoMa: A Dataset and Benchmark Towards Mobile Colorized Mapping Using a Backpack Multisensorial System
1 paper · 0 benchmarks
A dataset of 750 polymer abstracts annotated with the entity types: POLYMER, POLYMERCLASS, PROPERTYVALUE, PROPERTYNAME, MONOMER, ORGANICMATERIAL, INORGANICMATERIAL, and MATERIALAMOUNT.
1 paper · 0 benchmarks
This dataset was built with data acquired at the Hospital Clinic of Barcelona, Spain.
1 paper · 0 benchmarks
This dataset contains annual Sentinel-2 MSI composites (wet and dry season) for Kigali for the period 2016-2020.
1 paper · 0 benchmarks
A portrait dataset of images collected from Flickr.
1 paper · 0 benchmarks
https://drive.google.com/drive/folders/1ykFXI9AJjrRCmz7MvjdqdCq7e-4Hir-c This dataset provides several scenarios of imaging sonar images, supplemented with ground truth sensor pose.
1 paper · 0 benchmarks
Includes 36 models trained to role-play the following personas: - saint - truthteller - genie - moneymaximizer - fitnessmaximizer - rewardmaximizer The following is an example prompt: >You are an AI system.
1 paper · 0 benchmarks
This data is derived from the 7Scenes dataset.
1 paper · 0 benchmarks
Post-hoc Calibration Dataset This dataset collection is designed for the evaluation and development of post-hoc calibration methods for deep neural network classifiers.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
A dataset for studying situated goal-directed human communication.
1 paper · 0 benchmarks
In this Pre-Contest Workshop Slidedeck.pdf: Instructional materials delivered for the seven pre-contest workshops
1 paper · 0 benchmarks
In this Pre-Contest Workshop Video Recordings folder: Seven screen and audio recordings of seven pre-contest workshops
1 paper · 0 benchmarks
This repository contains ready-to-use frequency time series as well as the corresponding pre-processing scripts in python.
1 paper · 0 benchmarks
We release various types of word embeddings for multiple Indian languages.
1 paper · 0 benchmarks
PreRAID (Prescreening Rheumatoid Arthritis Information Database (PreRAID))
PreRAID is a structured dataset designed to evaluate the diagnostic capabilities of Large Language Models (LLMs) in Rheumatoid Arthritis (RA) diagnosis.
1 paper · 0 benchmarks
A pretrained PredNet neural network, used in EIGen to generate color illusions.
1 paper · 0 benchmarks
A pretrained PredNet neural network, used in EIGen to generate grayscale illusions.
1 paper · 0 benchmarks
Dataset Description High-level explanation of dataset characteristics: This dataset includes electromyographic (EMG) signals captured using the BiTalino device.
1 paper · 0 benchmarks
Press Briefing Claim Dataset The dataset contains a total of 53 press briefings from a time span of over four years (2017-2021).
1 paper · 0 benchmarks
A synthetic dataset with 206K pressure images with 3D human poses and shapes.
1 paper · 0 benchmarks
The pretrained models from four image translation algorithms: ACL-GAN, Council-GAN, CycleGAN, and U-GAT-IT on three benchmarking datasets: Selfie2Anime, CelebAgender, CelebAglasses.
1 paper · 0 benchmarks
A large-scale dataset for proactive document retrieval that consists of over 2.8 million conversations from Reddit.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.