Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 104 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 4945–4992 of 12,172
PDS-COCO (Photometrically Distorted Synthetic COCO)
Photometrically Distorted Synthetic COCO (PDS-COCO) dataset is a synthetically created dataset for homography estimation learning.
4 papers · 1 benchmark
PELD is a text-based emotional dialog dataset with personality traits for speakers.
4 papers · 0 benchmarks
PHD² (Personalized Highlight Detection Dataset)
The dataset contains information on what video segments a specific user considers a highlight.
4 papers · 0 benchmarks
PRO-teXt is an extension of PROXD with the inclusion of text prompts to synthesize objects.
4 papers · 2 benchmarks
PanLex-BLI (PanLex-based bilingual lexicons for 210 language pairs)
PanLex-based bilingual lexicons for 210 language pairs
4 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
4 papers · 1 benchmark
ParaPat (Parallel Corpus of Patents Abstracts)
A parallel corpus from the open access Google Patents dataset in 74 language pairs, comprising more than 68 million sentences and 800 million tokens.
4 papers · 0 benchmarks
ParaShoot is the first question answering dataset in modern Hebrew.
4 papers · 0 benchmarks
PeopleSansPeople (PeopleSansPeople: A Synthetic Data Generator for Human-Centric Computer Vision)
In recent years, person detection and human pose estimation have made great strides, helped by large-scale labeled datasets.
4 papers · 0 benchmarks
Persian dataset for relation extraction, which is an expert-translated version of the "Semeval-2010-Task-8" dataset.
4 papers · 0 benchmarks
Phy-Q is a benchmark that requires an agent to reason about physical scenarios and take an action accordingly.
4 papers · 0 benchmarks
PhyAAt (Physiology of Auditory Attention)
The dataset contains a collection of physiological signals (EEG, GSR, PPG) obtained from an experiment of the auditory attention on natural speech.
4 papers · 4 benchmarks
Data for this challenge were contributed by the Massachusetts General Hospital’s (MGH) Computational Clinical Neurophysiology Laboratory (CCNL), and the Clinical Data Animation Laboratory (CDAC).
4 papers · 2 benchmarks
Data The data for this Challenge are from multiple sources: CPSC Database and CPSC-Extra Database INCART Database PTB and PTB-XL Database The Georgia 12-lead ECG Challenge (G12EC) Database Undisclosed Database The first source is the…
4 papers · 1 benchmark
The Polaris dataset offers a large-scale, diverse benchmark for evaluating metrics for image captioning, surpassing existing datasets in terms of size, caption diversity, number of human judgments, and granularity of the evaluations.
4 papers · 0 benchmarks
ProofNet# is an evaluation benchmark derived from the original ProofNet, which contains 371 paired examples of informal undergraduate mathematical statements and their corresponding formalizations.
4 papers · 0 benchmarks
PsyMo (PsyMo: A Dataset for Estimating Self-Reported Psychological Traits from Gait)
Psychological trait estimation from external factors such as movement and appearance is a challenging and long-standing problem in psychology, and is principally based on the psychological theory of embodiment.
4 papers · 0 benchmarks
QC-Science contains 47832 question-answer pairs belonging to the science domain tagged with labels of the form subject - chapter - topic.
4 papers · 1 benchmark
QT-NSTDB (QT database + MIT-BIH Noise Stress Test Database (NSTDB))
We designed a baseline wander (BLW) removal benchmark to evaluate various methods using a consistent test set and uniform conditions.
4 papers · 1 benchmark
QUVA Repetition dataset consists of 100 videos displaying a wide variety of repetitive video dynamics, including swimming, stirring, cutting, combing and music-making.
4 papers · 0 benchmarks
Aims to help V-NLIs recognize analytic tasks from free-form natural language by training and evaluating cutting-edge multi-label classification models.
4 papers · 0 benchmarks
Collects dense per-video-shot concept annotations.
4 papers · 1 benchmark
Relatedness judgments of ambiguous English words, in experimentally controlled sentential contexts.
4 papers · 0 benchmarks
The Road Damage Dataset 2020 (RDD-2020) Secondly is a large-scale heterogeneous dataset comprising 26620 images collected from multiple countries using smartphones.
4 papers · 0 benchmarks
Real-M is a crowd-sourced speech-separation corpus of real-life mixtures.
4 papers · 0 benchmarks
RED (Real Embodied Dataset)
The Real Embodied Dataset (RED) is a computer vision large-scale dataset for grasping in cluttered scenes.
4 papers · 0 benchmarks
RISeC (Recipe Instruction Semantics Corpus)
We propose a newly annotated dataset for information extraction on recipes.
4 papers · 0 benchmarks
Deep neural networks for video based eye tracking have demonstrated resilience to noisy environments, stray reflections and low resolution.
4 papers · 0 benchmarks
RMAS (Real-World Marine Animal Segmentation)
We construct a new large-scale real-world MAS data set for conducting extensive experiments.
4 papers · 1 benchmark
RONEC (Romanian Named Entity Corpus)
Romanian Named Entity Corpus is a named entity corpus for the Romanian language.
4 papers · 0 benchmarks
Review-Rebuttal (RR) dataset is introduced to facilitate the study of argument pair extraction in the peer review and rebuttal domain.
4 papers · 1 benchmark
RTASC (ROBIN Technical Acquisition Speech Corpus)
The ROBIN Technical Acquisition Speech Corpus (ROBINTASC) was developed within the ROBIN project.
4 papers · 0 benchmarks
RWWD (Real World Worry Dataset)
Real World Worry Dataset (RWWD) captures the emotional responses of UK residents to COVID-19 at a point in time where the impact of the COVID19 situation affected the lives of all individuals in the UK.
4 papers · 0 benchmarks
We introduce a new database of voice recordings with the goal of supporting research on vulnerabilities and protection of voice-controlled systems.
4 papers · 0 benchmarks
RealMCVSR (Real-world Multi-Camera Video Super-Resolution)
Our RealMCVSR dataset provides real-world HD video triplets concurrently recorded by Apple iPhone 12 Pro Max equipped with triple cameras having fixed focal lengths: ultra-wide (30mm), wide-angle (59mm), and telephoto (147mm).
4 papers · 1 benchmark
RegDB-C is an evaluation set that consists of algorithmically generated corruptions applied to the RegDB test-set (color images).
4 papers · 0 benchmarks
Relative Human (RH) contains multi-person in-the-wild RGB images with rich human annotations, including: Depth layers: relative depth relationship/ordering between all people in the image.
4 papers · 2 benchmarks
The Restaurant-ACOS dataset is constructed based on the SemEval 2016 Restaurant dataset (Pontiki et al., 2016) and its expansion datasets (Fan et al., 2019; Xu et al., 2020).
4 papers · 1 benchmark
RetVQA (Retrieval-Based Visual Question Answering)
The RetVQA dataset is a large-scale dataset designed for Retrieval-Based Visual Question Answering (RetVQA).
4 papers · 1 benchmark
Real-world omnidirectional multi-view image dataset.
4 papers · 0 benchmarks
RidgeBase (RidgeBase: A Cross-Sensor Multi-Finger Contactless Fingerprint Dataset)
Contactless fingerprint matching using smartphone cameras can alleviate major challenges of traditional fingerprint systems including hygienic acquisition, portability and presentation attacks.
4 papers · 0 benchmarks
Text and signs around roads provide crucial information for drivers, vital for safe navigation and situational awareness.
4 papers · 1 benchmark
A dataset for robustness analysis of point cloud classification models (independent of data augmentation) to input transformations.
4 papers · 0 benchmarks
S2TLD (SJTU Small Traffic Light Dataset)
S2TLD is a traffic light dataset, which contains 5,786 images of approximately 1,080 1,920 pixels and 720 1,280 pixels.
4 papers · 0 benchmarks
S3E is a novel large-scale multimodal dataset captured by a fleet of unmanned ground vehicles along four designed collaborative trajectory paradigms.
4 papers · 0 benchmarks
SCDB (Simple Concept DataBase)
Includes annotations for 10 distinguishable concepts.
4 papers · 0 benchmarks
SCDE is a human-created sentence cloze dataset, collected from public school English examinations in China.
4 papers · 1 benchmark
The Situated Corpus Of Understanding Transactions (SCOUT) is a multi-modal collection of human-robot dialogue in the task domain of collaborative exploration.
4 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.