Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 57 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 2689–2736 of 3,998
GroundCap is a novel grounded image captioning dataset derived from MovieNet, containing 52,350 movie frames with detailed grounded captions.
1 paper · 0 benchmarks
For each problem, we provide 4 variants of prompts: 1.
1 paper · 0 benchmarks
HA-ViD (HA-ViD: A Human Assembly Video Dataset)
Understanding comprehensive assembly knowledge from videos is critical for futuristic ultra-intelligent industry.
1 paper · 0 benchmarks
HALvest is a textual dataset comprising 17 billion tokens in 56 languages and 13 domains.
1 paper · 0 benchmarks
HALvest-Geometric is a subset of HALvest: an academic citation network with 238,397 disambiguated authors and 18,662,037 scholarly papers.
1 paper · 0 benchmarks
HASCD (Human Activity Segmentation Challenge Dataset)
HASCD (Human Activity Segmentation Challenge Dataset) contains 250 annotated multivariate time series capturing 10.7 h of real-world human motion smartphone sensor data from 15 bachelor computer science students.
1 paper · 0 benchmarks
HAVOC (Harmful Abstractions and Violations in Open Completions Benchmark)
measure the toxicity generated by language models across input severity and harm categories, by creating a new benchmark of open ended prefixes.
1 paper · 0 benchmarks
The HDRT dataset is a large-scale dataset designed for infrared-guided high dynamic range (HDR) imaging.
1 paper · 0 benchmarks
HDT-QA (human driving test question answering dataset)
HDT-QA, coupled with driving manuals, offers an extensive compendium of driving instructions and driving knowledge tests across all 51 states of the US.
1 paper · 0 benchmarks
HEIMT (Hyperscanning EEG of Interactive Math Task)
Continuous EEG activity was recorded from each member of the dyad using an ActiveTwo head cap and the ActiveTwo Biosemi system (BioSemi, Amsterdam, Netherlands).
1 paper · 0 benchmarks
HGP (Hands Guns and Phones Dataset)
Hands Guns and Phones (HGP) dataset contains 2199 images (1989 for training an 210 for testing) of people using guns or phones in real-world scenarios (people making phones reviews, shooting drills, or making calls).
1 paper · 0 benchmarks
HOWS-CL-25 (Household Objects Within Simulation dataset for Continual Learning) is a synthetic dataset especially designed for object classification on mobile robots operating in a changing environment (like a household), where it is…
1 paper · 2 benchmarks
HPO (Human Phenotype Ontology)
The Human Phenotype Ontology (HPO) graph is a standardized vocabulary of human phenotypic abnormalities and their relationships.
1 paper · 0 benchmarks
HRI (High-resolution Rainy Image)
The HRI Dataset comprises a total of 3,200 image pairs.
1 paper · 0 benchmarks
The dataset concerns toy tasks that a human should teach to a robot.
1 paper · 0 benchmarks
HRPlanesV2 (HRPlanesv2 - High Resolution Satellite Imagery for Aircraft Detection)
The HRPlanesv2 dataset contains 2120 VHR Google Earth images.
1 paper · 0 benchmarks
HSIRS (High-quality Spectral Image Resonstruction and Segmentation Dataset)
We introduce HSIRS, a large scale dataset of hyper-spectral images along with corresponding manually annotated segmentation maps for material characterization and classification based on spectral signature.
1 paper · 0 benchmarks
Human fibrosarcoma HT1080WT (ATCC) cells at low cell densities embedded in 3D collagen type I matrices [1].
1 paper · 0 benchmarks
The dataset comprises 2886 patches in total (2 m GSD), of which 1732 patches for training and 1154 patches for testing.
1 paper · 1 benchmark
HaSPeR (Hand Shadow Puppet Image Repository)
TODO
1 paper · 0 benchmarks
Hand Wash Dataset consists of 292 videos of hand washes with each hand wash having 12 steps, for a total of 3,504 clips, in different environments to provide as much variance as possible.
1 paper · 0 benchmarks
HarmfulTasks (Harmful and Malicious Tasks for LLMs in Jailbreaking Prompts)
This dataset consists of 225 malicious tasks, which were integrated into ten distinct jailbreaking prompts.
1 paper · 0 benchmarks
The National Health and Nutrition Examination Survey (NHANES) provides data on the health and environmental exposure of the non-institutionalized US population.
1 paper · 0 benchmarks
Harry Potter Dialogue is the first dialogue dataset that integrates with scene, attributes and relations which are dynamically changed as the storyline goes on.
1 paper · 2 benchmarks
Multi-Modal Hate Speech Detection with Graph Context.
1 paper · 0 benchmarks
Hawk Annotation Dataset includes language descriptions specifically for anomaly scenes in seven existing video anomaly datasets.
1 paper · 0 benchmarks
Heritage Pointcloud Instance Collection dataset, acquired from two large buildings and annotated at a point-wise semantic level based on existent BIM models.
1 paper · 1 benchmark
Inpatient claims, Outpatient claims and Beneficiary details of each provider.
1 paper · 1 benchmark
See https://zenodo.org/record/5500215#.YUCgD51Kg2w
1 paper · 0 benchmarks
The Helsinki Prosody Corpus is a dataset for predicting prosodic prominence from written text.
1 paper · 1 benchmark
Each episode directory contains word-level and segment-level information of the whole episode and also parallel samples extracted under segmentseng and segmentsspa subdirectories.
1 paper · 0 benchmarks
Heteroatom doped graphene supercapacitor feature data is gathered from various literatures for use in machine learning tasks.
1 paper · 0 benchmarks
Hinglish-TOP is a human annotated code-switched semantic parsing dataset containing 10k human annotations for Hindi-English (HINGLISH) code switched utterances, and over 170K CST5 generated code-switched utterances from the TOPv2 dataset.
1 paper · 0 benchmarks
This dataset is composed of 7,753 pairs of whole slide images and their corresponding diagnostic reports, extracted from the TCGA platform and refined with large language models.
1 paper · 1 benchmark
Here the dataset described in Hitchhiking Rides Dataset: Two decades of crowd-sourced records on stochastic traveling(https://arxiv.org/abs/2506.21946) is published.
1 paper · 0 benchmarks
HpVaxFrames includes 64 Vaccine Hesitancy Framings found on Twitter about the HPV vaccines.
1 paper · 0 benchmarks
HuSHeM (Human Sperm Head Morphology Dataset)
At the Isfahan Fertility and Infertility Center, semen samples were collected from fifteen patients.
1 paper · 0 benchmarks
A dataset including texts by humans (labeled 0) and then rephrased by ChatGPT (labeled 1), created to train models for machine-generated text detection.
1 paper · 0 benchmarks
HumanMT is a collection of human ratings and corrections of machine translations.
1 paper · 0 benchmarks
The image collection of the IAPR TC-12 Benchmark consists of 20,000 still natural images taken from locations around the world and comprising an assorted cross-section of still natural images.
1 paper · 0 benchmarks
The IAW dataset contains 420 Ikea furniture pieces from 14 common categories e.g.
1 paper · 0 benchmarks
This dataset contains general and named entities annotations on both clean written text and on noisy speech data.
1 paper · 0 benchmarks
The dataset is composed of Hematoxylin and eosin (H&E) stained breast histology microscopy and whole-slide images.
1 paper · 1 benchmark
A maintained database tracks ICLR submissions and reviews, augmented with author profiles and higher-level textual features.
1 paper · 0 benchmarks
The dataset is taken from the First shared task on Information Extractor for Conversational Systems in Indian Languages (IECSIL) .
1 paper · 1 benchmark
IEIs (Ion and Electron Insulators)
We would like to introduce three types of ion and electron insulators, i.e.
1 paper · 0 benchmarks
IHEval (Evaluation on Instruction Hierarchy)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.