Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 181 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 8641–8688 of 12,172
HFFD (Hybrid Fake Face Dataset)
We build a hybrid fake face (HFF) dataset, which contains eight types of face images.
1 paper · 0 benchmarks
HGP (Hands Guns and Phones Dataset)
Hands Guns and Phones (HGP) dataset contains 2199 images (1989 for training an 210 for testing) of people using guns or phones in real-world scenarios (people making phones reviews, shooting drills, or making calls).
1 paper · 0 benchmarks
The HH Red Teaming dataset comprises two distinct types of data, each serving a unique purpose: 1.
1 paper · 0 benchmarks
HICRD (Heron Island Coral Reef Dataset)
HICRD (Heron Island Coral Reef Dataset) is a large-scale real underwater image dataset for underwater image restoration.
1 paper · 0 benchmarks
HIU-DMTL-Data (Hand Image Understanding via Deep Multi-Task Learning.)
See the paper for more details.
1 paper · 0 benchmarks
Habitat-Matterport 3D Semantics Dataset (HM3D-Semantics v0.1) is the largest-ever dataset of semantically-annotated 3D indoor spaces.
1 paper · 0 benchmarks
This dataset contains more than 700,000 unique voltage vs.
1 paper · 0 benchmarks
A dataset for pose estimation of hand when interacting with object and severe occlusions.
1 paper · 0 benchmarks
HOI-SDC (Setting for Double Challenge of Human-Object Interaction Detection)
In order to avoid the training process of the model being influenced by a portion of HOI classes with a very small number of instances, we remove some of the HOI classes containing a very small number of instances and HOI classes with no…
1 paper · 0 benchmarks
HOWS-CL-25 (Household Objects Within Simulation dataset for Continual Learning) is a synthetic dataset especially designed for object classification on mobile robots operating in a changing environment (like a household), where it is…
1 paper · 2 benchmarks
HPO (Human Phenotype Ontology)
The Human Phenotype Ontology (HPO) graph is a standardized vocabulary of human phenotypic abnormalities and their relationships.
1 paper · 0 benchmarks
HPointLoc is a dataset designed for exploring capabilities of visual place recognition in indoor environment and loop detection in simultaneous localization and mapping.
1 paper · 0 benchmarks
Dataset Card HR-Multiwoz is a fully-labeled dataset of 5980 extractive qa spanning 10 HR domains to evaluate LLM Agent.
1 paper · 0 benchmarks
HR-Crime is a subset of the UCF-Crime dataset suitable for human-related anomaly detection tasks.
1 paper · 0 benchmarks
HR-Multiwoz is a fully-labeled dataset of 550 conversations spanning 10 HR domains to evaluate LLM Agent.
1 paper · 0 benchmarks
HRI (High-resolution Rainy Image)
The HRI Dataset comprises a total of 3,200 image pairs.
1 paper · 0 benchmarks
The dataset concerns toy tasks that a human should teach to a robot.
1 paper · 0 benchmarks
HRPlanesV2 (HRPlanesv2 - High Resolution Satellite Imagery for Aircraft Detection)
The HRPlanesv2 dataset contains 2120 VHR Google Earth images.
1 paper · 0 benchmarks
HS-BAN is a binary class hate speech (HS) dataset in Bangla language consisting of more than 50,000 labeled comments, including 40.17% hate and rest are non hate speech.
1 paper · 0 benchmarks
HSIRS (High-quality Spectral Image Resonstruction and Segmentation Dataset)
We introduce HSIRS, a large scale dataset of hyper-spectral images along with corresponding manually annotated segmentation maps for material characterization and classification based on spectral signature.
1 paper · 0 benchmarks
HT Docking is a dataset consisting of 200 million 3D complex structures and 2D structure scores across a consistent set of 13 million in-stock'' molecules over 15 receptors, or binding sites, across the SARS-CoV-2 proteome.
1 paper · 0 benchmarks
Human fibrosarcoma HT1080WT (ATCC) cells at low cell densities embedded in 3D collagen type I matrices [1].
1 paper · 0 benchmarks
HTDM (Hypertention Disease Medication)
Hypertention Disease Medication dataset.
1 paper · 0 benchmarks
We propose a new benchmark called Human Video Instance Segmentation (HVIS), which focuses on complex real-world scenarios with sufficient human instance masks and identities.
1 paper · 0 benchmarks
HW-NAS-Bench is a dataset for HardWare-aware Neural Architecture Search (HW-NAS).
1 paper · 0 benchmarks
HYPE (PPG and Blood Pressure from a Hypertensive Population)
HYPE Dataset - Version 1.0.0 REFERENCE PAPER ------------------- Morassi Sasso, A., Datta, S., Jeitler, M., Steckhan, N., Kessler, C.
1 paper · 0 benchmarks
The dataset comprises 2886 patches in total (2 m GSD), of which 1732 patches for training and 1154 patches for testing.
1 paper · 1 benchmark
HYouTube is a video for Video harmonization, which aims to adjust the foreground of a composite video to make it compatible with the background.
1 paper · 0 benchmarks
HaSPeR (Hand Shadow Puppet Image Repository)
TODO
1 paper · 0 benchmarks
HabiCrowd, a new dataset and benchmark for crowd-aware visual navigation that surpasses other benchmarks in terms of human diversity and computational utilization.
1 paper · 0 benchmarks
Halpe-FullBody is a full body keypoints dataset where each person has annotated 136 keypoints, including 20 for body, 6 for feet, 42 for hands and 68 for face.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This is a dataset for benchmarking in-hand manipulation on different robot platforms.
1 paper · 0 benchmarks
Hand Wash Dataset consists of 292 videos of hand washes with each hand wash having 12 steps, for a total of 3,504 clips, in different environments to provide as much variance as possible.
1 paper · 0 benchmarks
This is an image database of Handwritten Devanagari characters.
1 paper · 0 benchmarks
The HardZiPA folder contains illuminance and RGB data as well as CO2 and TVOC data for five sensing devices.
1 paper · 0 benchmarks
HarmfulTasks (Harmful and Malicious Tasks for LLMs in Jailbreaking Prompts)
This dataset consists of 225 malicious tasks, which were integrated into ten distinct jailbreaking prompts.
1 paper · 0 benchmarks
Beats, downbeats, and functional structural annotations for 912 Pop tracks.
1 paper · 2 benchmarks
The National Health and Nutrition Examination Survey (NHANES) provides data on the health and environmental exposure of the non-institutionalized US population.
1 paper · 0 benchmarks
Harry Potter Dialogue is the first dialogue dataset that integrates with scene, attributes and relations which are dynamically changed as the storyline goes on.
1 paper · 2 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
HateSpeechCorpus is a dataset collected from Twitter, consisting of 3003 tweets.
1 paper · 0 benchmarks
Multi-Modal Hate Speech Detection with Graph Context.
1 paper · 0 benchmarks
Hawk Annotation Dataset includes language descriptions specifically for anomaly scenes in seven existing video anomaly datasets.
1 paper · 0 benchmarks
The Haydn Annotation Dataset consists of note onset annotations from 24 experiment participants with varying musical experience.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.