Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 204 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 9745–9792 of 12,172
This dataset was collected by a collaboration of researchers from Children’s Wisconsin, Marquette University, Varian Medical Systems, Medical College of Wisconsin, and Stanford University as part of a project funded by the National…
1 paper · 0 benchmarks
Data-set from "PEEK-An LSTM Recurrent Network for Motion Classification from Sparse Data"
1 paper · 0 benchmarks
Peer to Peer Hate is a comprehensive hate speech dataset capturing various types of hate.
1 paper · 0 benchmarks
The dataset is derived from the MSK-IMPACT dataset designed and published by Zehir using the code published by Penson et al..
1 paper · 0 benchmarks
Pentachromatic Cultural Palette Dataset is characterized by unique cultural semantics and values.
1 paper · 0 benchmarks
PerPaDa is a Persian paraphrase dataset that is collected from users' input in a plagiarism detection system.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
PIE stands for Performance Improving Code Edits.
1 paper · 0 benchmarks
The Perfume Co-Preference Network dataset comprises comprehensive user reviews and ratings collected from the Persian retail platform Atrafshan.
1 paper · 0 benchmarks
Classification dataset of 90 molecules according to its effects in the circadian rhythm from the UCI Machine Learning Repository.
1 paper · 0 benchmarks
The Peripheral Blood Cell} (PBC) dataset consists of 17,092 images.
1 paper · 0 benchmarks
This dataset contains the results of a depression screening experiment using two instruments: The PHQ-9 depression screening questionnaire and the chabot Perla.
1 paper · 0 benchmarks
The Permuted bAbi dialog task is an adaptation of the "Dialog bAbI tasks data" dataset released by Facebook.
1 paper · 0 benchmarks
Persian Font Recognition (PFR) A dataset in order to solve font recognition for the Persian language.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Data collection was conducted by asking some adults from social media and some students from an elementary school to participate in our experiment.
1 paper · 0 benchmarks
The Persian Reverse Dictionary Dataset is a collection of 855217 words along with the phrases describing them.
1 paper · 0 benchmarks
Persian Text Image Segmentation (PTI SEG) This dataset is part of a paper titled "Persis: A Persian Font Recognition Pipeline Using Convolutional Neural Networks".
1 paper · 1 benchmark
PersianQA (Persian Question Answering Dataset)
PersianQA: a dataset for Persian Question Answering Persian Question Answering (PersianQA) Dataset is a reading comprehension dataset on Persian Wikipedia.
1 paper · 0 benchmarks
Persistence Diagram Benchmark
1 paper · 0 benchmarks
The PEDC is a corpus of 14 episodes of This American Life podcast transcripts that have been annotated for events.
1 paper · 0 benchmarks
Persuasive Writing Strategy dataset on the health subset of the Multi-FC dataset.
1 paper · 0 benchmarks
Pesteh-Set is made of two parts.
1 paper · 0 benchmarks
Briefly describe the dataset.
1 paper · 0 benchmarks
PhD (PhD: A ChatGPT-Prompted Visual hallucination Evaluation Dataset)
Multimodal Large Language Models (MLLMs) hallucinate, resulting in an emerging topic of visual hallucination evaluation (VHE).
1 paper · 0 benchmarks
This dataset is part of the journal paper "Discipline Reputation Evaluation Based on PhD Exchange Network", Author: Shudong YANG @dalian University of Technology.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
PheMT is a phenomenon-wise dataset designed for evaluating the robustness of Japanese-English machine translation systems.
1 paper · 0 benchmarks
PhilPapers is a remarkable resource for the philosophical community.
1 paper · 1 benchmark
We employ a nationwide phone call dataset from Jan.
1 paper · 0 benchmarks
PhotoMatte85 contains 85 protrait images.
1 paper · 0 benchmarks
Trained models, evaluation data sets and results, and references to the full data used in the training, validation and testing of the models
1 paper · 0 benchmarks
A large-scale dataset of user annotations on seven common photographic defects.
1 paper · 0 benchmarks
Photozilla is a large-scale dataset which includes over 990k images belonging to 10 different photographic styles.
1 paper · 0 benchmarks
Phrase in Context is a curated benchmark for phrase understanding and semantic search, consisting of three tasks of increasing difficulty: Phrase Similarity (PS), Phrase Retrieval (PR) and Phrase Sense Disambiguation (PSD).
1 paper · 0 benchmarks
PhysNLU is a collection of 4 core datasets related to sentence classification, ordering, and coherence of physics explanations based on related tasks.
1 paper · 0 benchmarks
Physical concept understanding benchmark.
1 paper · 0 benchmarks
PicTropes is a dataset of films and the tropes that they use created from the database DBTropes.org.
1 paper · 0 benchmarks
Pick-a-Filter is a semi-synthetic dataset constructed from Pick-a-Pic v1 to measure the capability of text-to-image models of adapting to heterogeneous preferences.
1 paper · 0 benchmarks
Deep learning for site safety: Real-time detection of personal protective equipment
1 paper · 0 benchmarks
This images has been collected from Pinterest and cropped.
1 paper · 0 benchmarks
The Pinterest Complete the Look dataset consists of over 1 million outfits and 4 million objects.
1 paper · 0 benchmarks
Pirá (Pirá: A Bilingual Portuguese-English Dataset for Question-Answering about the Ocean)
A large set of questions and answers about the ocean and the Brazilian coast both in Portuguese and English.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
an image cover dataset in short video recommendation
1 paper · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.