Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 182 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 8689–8736 of 12,172
Contrary to prior scene graph datasets, Haystack contains explicit negative annotations, i.e.
1 paper · 0 benchmarks
Heritage Pointcloud Instance Collection dataset, acquired from two large buildings and annotated at a point-wise semantic level based on existent BIM models.
1 paper · 1 benchmark
HeSum (Abstractive Text Summarization in Hebrew)
While large language models (LLMs) excel in various natural language tasks in English, their performance in lower-resourced languages like Hebrew, especially for generative tasks such as abstractive summarization, remains unclear.
1 paper · 0 benchmarks
Inpatient claims, Outpatient claims and Beneficiary details of each provider.
1 paper · 1 benchmark
Prediction of a speaker's height is of interest in fields such as voice forensics, surveillance, and automatic speaker profiling.
1 paper · 0 benchmarks
See https://zenodo.org/record/5500215#.YUCgD51Kg2w
1 paper · 0 benchmarks
Hello Watt (Hello Watt electricity consumption curves)
Hello Watt collects power usage data at a resolution of 30 minutes.
1 paper · 0 benchmarks
HelloWorld is a dataset of kinesthetic demonstrations collected using a Franka Emika Panda robot.
1 paper · 0 benchmarks
The Helsinki Prosody Corpus is a dataset for predicting prosodic prominence from written text.
1 paper · 1 benchmark
HengamCopus is a Persian corpus with temporal tags (BIO standard tagging scheme).
1 paper · 1 benchmark
Herbarium 2022 (Identify plant species of the Americas from herbarium specimens)
The Herbarium 2022: Flora of North America is a part of a project of the New York Botanical Garden funded by the National Science Foundation to build tools to identify novel plant species around the world.
1 paper · 1 benchmark
The FGVC 2019 Herbarium Challenge is to identify melastome species from herbarium specimens provided by the New York Botanical Garden (NYBG).
1 paper · 0 benchmarks
Each episode directory contains word-level and segment-level information of the whole episode and also parallel samples extracted under segmentseng and segmentsspa subdirectories.
1 paper · 0 benchmarks
Heteroatom doped graphene supercapacitor feature data is gathered from various literatures for use in machine learning tasks.
1 paper · 0 benchmarks
Hi-Phy is a benchmark for physical reasoning that allows researchers to test individual physical reasoning capabilities.
1 paper · 0 benchmarks
This dataset releases a significantly sized standard-abiding Hindi NER dataset containing 109,146 sentences and 2,220,856 tokens, annotated with 3 collapsed tags (PER, LOC, ORG).
1 paper · 1 benchmark
HiNER-original (HiNER: A Large Hindi Named Entity Recognition Dataset)
This dataset releases a significantly sized standard-abiding Hindi NER dataset containing 109,146 sentences and 2,220,856 tokens, annotated with 11 tags.
1 paper · 1 benchmark
For testing refusal behavior in a language-specific setting, we introduce HiXSTest — a set of manually curated prompts in the Hindi language designed to measure exaggerated safety.
1 paper · 0 benchmarks
HierText is the first dataset featuring hierarchical annotations of text in natural scenes and documents.
1 paper · 1 benchmark
The dataset contains 179 photographs taken by a UAV flying at 120 meters altitude.
1 paper · 0 benchmarks
This dataset contains stereo images, depth data, camera position/orientation, and camera information for 100 sorghum panicles (the seed-bearing head of the sorghum stalk), as well as semantic segmentation labels for a subset of the data.
1 paper · 0 benchmarks
Optimised constellation for the paper High-Cardinality Geometrical Constellation Shaping for the Nonlinear Fibre Channel.
1 paper · 0 benchmarks
This benchmark is based on the HILTI-OXFORD Dataset, which has been collected on construction sites as well as on the famous Sheldonian Theatre in Oxford, providing a large range of difficult problems for SLAM.
1 paper · 0 benchmarks
Hinglish-TOP is a human annotated code-switched semantic parsing dataset containing 10k human annotations for Hindi-English (HINGLISH) code switched utterances, and over 170K CST5 generated code-switched utterances from the TOPv2 dataset.
1 paper · 0 benchmarks
This dataset is composed of 7,753 pairs of whole slide images and their corresponding diagnostic reports, extracted from the TCGA platform and refined with large language models.
1 paper · 1 benchmark
Here the dataset described in Hitchhiking Rides Dataset: Two decades of crowd-sourced records on stochastic traveling(https://arxiv.org/abs/2506.21946) is published.
1 paper · 0 benchmarks
HoMG is a holoscopic 3D micro-gesture dataset captured with a holoscopic 3D camera.
1 paper · 0 benchmarks
HoaxItaly consists of over 1 million tweets shared during 2019 and containing links to thousands of news articles published on two classes of Italian outlets: (1) disinformation websites, i.e.
1 paper · 0 benchmarks
Dataset Description - Paper: TBC - Point of Contact: Josh McGiff (Josh.McGiff@ul.ie) Dataset Summary This dataset was developed to address the significant gap in online hate speech detection, particularly focusing on homophobia, which is…
1 paper · 0 benchmarks
The directory HiCIS contains two datasets for instance segmentation of honeycombs in concrete in COCO Format.
1 paper · 0 benchmarks
Collection of images of garbages grouped into 10 classes (metal, glass, biological, paper, battery, trash, cardboard, shoes, clothes, and plastic).
1 paper · 0 benchmarks
This dataset is used for predicting house prices from both images and textual information.
1 paper · 0 benchmarks
HowSumm is a large-scale query-focused multi-document summarization dataset.
1 paper · 2 benchmarks
HpVaxFrames includes 64 Vaccine Hesitancy Framings found on Twitter about the HPV vaccines.
1 paper · 0 benchmarks
HuSHeM (Human Sperm Head Morphology Dataset)
At the Isfahan Fertility and Infertility Center, semen samples were collected from fifteen patients.
1 paper · 0 benchmarks
The image set contains 180 high-resolution color microscopic images of human duodenum adenocarcinoma HuTu 80 cell populations obtained in an in vitro scratch assay (for the details of the experimental protocol, we refer to (Liang et al.,…
1 paper · 1 benchmark
Huebner2017 MOABB (Learning from label proportions for a visual matrix speller (ERP) dataset from Hübner et al 2017.)
1 paper · 1 benchmark
Huebner2018 MOABB (Mixture of LLP and EM for a visual matrix speller (ERP) dataset from Hübner et al 2018.)
1 paper · 1 benchmark
A dataset including texts by humans (labeled 0) and then rephrased by ChatGPT (labeled 1), created to train models for machine-generated text detection.
1 paper · 0 benchmarks
Human-Robot Interaction Conversational User Enjoyment Scale (HRI CUES) and this corresponding dataset aim to provide tools for measuring user enjoyment from an external perspective to supplement self-reported user enjoyment responses in…
1 paper · 0 benchmarks
Dataset used in Bütepage, Judith, et al.
1 paper · 0 benchmarks
HumanMT is a collection of human ratings and corrections of machine translations.
1 paper · 0 benchmarks
Language models (LMs) as conversational assistants recently became popular tools that help people accomplish a variety of tasks.
1 paper · 0 benchmarks
Overview - Dataset Name: HumanRig - Paper: CVPR2025 - "HumanRig: Learning Automatic Rigging for Humanoid Character in a Large Scale Dataset" - Authors: [Zedong Chu · Feng Xiong · Meiduo Liu · Jinzhi Zhang · Mingqi Shao · Zhaoxu Sun · Di…
1 paper · 0 benchmarks
The HumanoidRobotPose dataset is a dataset for real-time pose estimation of humanoid robots.
1 paper · 0 benchmarks
Hyperbard is a dataset of diverse relational data representations derived from Shakespeare's plays.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.