Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 85 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 4033–4080 of 12,172
The PRONOSTIA (also called FEMTO) bearing dataset consists of 17 accelerated run-to-failures on a small bearing test rig.
6 papers · 0 benchmarks
Starting from the Panoptic Dataset, we use the PanopTOP framework to generate the PanopTOP31K dataset, consisting of 31K images from 23 different subjects recorded from diverse and challenging viewpoints, also including the top-view.
6 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
6 papers · 1 benchmark
Peir Gross (Jing et al., 2018) was collected with descriptions in the Gross sub-collection from PEIR digital library, resulting in 7.442 image-caption pairs from 21 different sub-categories.
6 papers · 1 benchmark
Object pose estimation is crucial for robotic applications and augmented reality.
6 papers · 0 benchmarks
PhyBench is a comprehensive Text-to-Image (T2I) evaluation dataset designed to assess the physical commonsense of T2I models¹.
6 papers · 0 benchmarks
Data Description The training data contains twelve-lead ECGs.
6 papers · 2 benchmarks
The PieAPP dataset is a large-scale dataset used for training and testing perceptually-consistent image-error prediction algorithms.
6 papers · 0 benchmarks
The dataset is split between train, test and val folders.
6 papers · 0 benchmarks
This is a dataset of over 333K Pull Requests, used for automatic pull request description generation.
6 papers · 0 benchmarks
Q-Traffic is a large-scale traffic prediction dataset, which consists of three sub-datasets: query sub-dataset, traffic speed sub-dataset and road network sub-dataset.
6 papers · 1 benchmark
QA-SRL Bank 2.0 is a large-scale corpus of Question-Answer driven Semantic Role Labeling (QA-SRL) annotations.
6 papers · 0 benchmarks
QDax is a benchmark suite designed for for Deep Neuroevolution in Reinforcement Learning domains for robot control.
6 papers · 0 benchmarks
QM8 dataset is a collection of molecular data used for studying quantum mechanical calculations of electronic spectra and excited state energy of small molecules.
6 papers · 1 benchmark
The RAD-ChestCT dataset is a large medical imaging dataset developed by Duke MD/PhD Rachel Draelos during her Computer Science PhD supervised by Lawrence Carin.
6 papers · 0 benchmarks
For RAWFC, we constructed it from scratch by collecting the claims from Snopes and relevant raw reports by retrieving claim keywords.
6 papers · 1 benchmark
RCB (Russian Commitment Bank)
The Russian Commitment Bank is a corpus of naturally occurring discourses whose final sentence contains a clause-embedding predicate under an entailment cancelling operator (question, modal, negation, antecedent of conditional).
6 papers · 1 benchmark
This dataset arises from the READ project (Horizon 2020).
6 papers · 1 benchmark
REFIT (An electrical load measurements dataset of United Kingdom households from a two-year longitudinal study)
Smart meter roll-outs provide easy access to granular meter measurements, enabling advanced energy services, ranging from demand response measures, tailored energy feedback and smart home/building automation.
6 papers · 0 benchmarks
RELX is a benchmark dataset for cross-lingual relation classification in English, French, German, Spanish and Turkish.
6 papers · 0 benchmarks
Rad-ReStruct is a fine-grained structured reporting dataset for Chest X-Ray images.
6 papers · 0 benchmarks
10,000 news collected from a social network in Vietnam.
6 papers · 0 benchmarks
ReadingBank is a benchmark dataset for reading order detection built with weak supervision from WORD documents, which contains 500K document images with a wide range of document types as well as the corresponding reading order information.
6 papers · 1 benchmark
RealCQA Scientific Chart Question Answering as a Test-bed for First-Order Logic check on huggingface : https://huggingface.co/datasets/sal4ahm/RealCQA
6 papers · 1 benchmark
Part of the Controlled Noisy Web Labels Dataset.
6 papers · 2 benchmarks
Part of the Controlled Noisy Web Labels Dataset.
6 papers · 2 benchmarks
Part of the Controlled Noisy Web Labels Dataset.
6 papers · 2 benchmarks
Reddit Conversation Corpus (RCC) consists of conversations, scraped from Reddit, for a 20 month period from November 2016 until August 2018.
6 papers · 0 benchmarks
The Relative Size dataset contains 486 object pairs between 41 physical objects.
6 papers · 0 benchmarks
A dataset for text in driving videos.
6 papers · 0 benchmarks
We contribute a new challenging visual grounding dataset for robotic perception and reasoning in indoor environments, called RoboRefIt.
6 papers · 0 benchmarks
RxRx1 is a biological dataset designed specifically for the systematic study of batch effect correction methods.
6 papers · 1 benchmark
Includes 4405 images with 111251 heads annotated.
6 papers · 0 benchmarks
SDWPF (A Dataset for Spatial Dynamic Wind Power Forecasting Challenge at KDD Cup 2022)
The unique Spatial Dynamic Wind Power Forecasting dataset: SDWPF, which includes the spatial distribution of wind turbines, as well as the dynamic context factors.
6 papers · 0 benchmarks
Test set version 2 for the San Francisco eXtra Large dataset
6 papers · 1 benchmark
SMART-101 (Simple Multimodal Algorithmic Reasoning Task Dataset)
Recent times have witnessed an increasing number of applications of deep neural networks towards solving tasks that require superior cognitive abilities, e.g., playing Go, generating art, ChatGPT, etc.
6 papers · 0 benchmarks
This corpus has been collected from free or free for research sources at the Internet: - A collection of 425 SMS spam messages was manually extracted from the Grumbletext Web site.
6 papers · 1 benchmark
SPACE is a simulator for physical Interactions and causal learning in 3D environments.
6 papers · 0 benchmarks
A dataset of utterances, incorrect SQL interpretations and the corresponding natural language feedback.
6 papers · 0 benchmarks
SWIMSEG (Singapore Whole sky IMaging SEGmentation Database)
The SWIMSEG dataset contains 1013 images of sky/cloud patches, along with their corresponding binary segmentation maps.
6 papers · 1 benchmark
SWSR (Sina Weibo Sexism Review)
The Sina Weibo Sexism Review (SWSR) dataset is a dataset to research online sexism in Chinese.
6 papers · 0 benchmarks
Dataset with 625,000 ethical judgments over 32,000 real-life anecdotes.
6 papers · 0 benchmarks
SemEval 2014 is a collection of datasets used for the Semantic Evaluation (SemEval) workshop, an annual event that focuses on the evaluation and comparison of systems that can analyze diverse semantic phenomena in text.
6 papers · 0 benchmarks
A multimodal dataset for sentiment analysis on internet memes.
6 papers · 0 benchmarks
SemOpenAlex is an extensive RDF knowledge graph that contains over 26 billion triples about scientific publications and their associated entities, such as authors, institutions, journals, and concepts.
6 papers · 0 benchmarks
ShapeTalk contains over half a million discriminative utterances produced by contrasting the shapes of common 3D objects for a variety of object classes and degrees of similarity.
6 papers · 0 benchmarks
ShellcodeIA32 is a dataset containing 20 years of shellcodes from a variety of sources is the largest collection of shellcodes in assembly available to date.
6 papers · 1 benchmark
SherLIiC is a testbed for lexical inference in context (LIiC), consisting of 3985 manually annotated inference rule candidates (InfCands), accompanied by (i) ~960k unlabeled InfCands, and (ii) ~190k typed textual relations between Freebase…
6 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.