Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 40 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 1873–1920 of 3,998

We release the dataset for non-commercial research.
2 papers · 0 benchmarks
GQA-OOD is a new dataset and benchmark for the evaluation of VQA models in OOD (out of distribution) settings.
2 papers · 0 benchmarks
GVLQA (Graph Vision-Language Question-Answering)
GVLQA is the first vision-language QA dataset for general graph reasoning.
2 papers · 0 benchmarks
Global Wheat Head 2021 (Global Wheat Head Dataset 2021)
Global WHEAT Dataset 2021 is the extentions of the Global Wheat Dataset 2020.
2 papers · 0 benchmarks
The Gun Violence Corpus (GVC) consists of 241 unique incidents for which we have structured data on a) location, b) time c) the name, gender and age of the victims and d) the status of the victims after the incident: killed or injured.
2 papers · 0 benchmarks
Automated measurement of fetal head circumference using 2D ultrasound images
2 papers · 0 benchmarks
HCP Aging (Lifespan Human Connectome Project Aging)
Lifespan HCP Release 2.0 includes cross-sectional visit 1 (V1) preprocessed structural and functional imaging data, unprocessed V1 imaging data for all included modalities (structural, high-res hippocampal T2, resting state fMRI, task…
2 papers · 1 benchmark
HERA RFI Detection (Hydrogen Epoch of Reionization Array (HERA))
This dataset contains simulated and expert-labelled spectrograms from two radio telescopes: the Hydrogen Epoch of Reionization Array (HERA) in South Africa and the Low-Frequency Array (LOFAR) in the Netherlands.
2 papers · 1 benchmark
HERDPhobia is an annotated hate speech detection dataset on Fulani herders in Nigeria -- in three languages: English, Nigerian-Pidgin, and Hausa.
2 papers · 0 benchmarks
HLGD (Headline Grouping Dataset)
The Headline Grouping dataset is a binary classification dataset on pairs of news headline.
2 papers · 0 benchmarks
HalluEditBench is a comprehensive benchmark for evaluating knowledge editing methods' effectiveness in correcting real-world hallucinations.
2 papers · 0 benchmarks
This is a Twitter dataset of 100,386 users along with up to 200 tweets from their timelines with a random-walk-based crawler on the retweet graph, with a subsample of 4,972 which is manually annotated as hateful or not through…
2 papers · 0 benchmarks
HeriGraph (Multimodal Machine Learning Datasets on Graphs of Heritage Values and Attributes)
The dataset contains constructed multi-modal features (visual and textual), pseudo-labels (on heritage values and attributes), and graph structures (with temporal, social, and spatial links) constructed using User-Generated Content data…
2 papers · 0 benchmarks
HiAML Computational Graph (CG) family introduced in "GENNAPE: Towards Generalized Neural Architecture Performance Estimators", accepted to AAAI-23.
2 papers · 0 benchmarks
The Horne 2017 Fake News Data contains two independed news datasets: 1.
2 papers · 0 benchmarks
Hotel (Hospitality > Tourism > Hotel Demand/Sales)
The dataset contains the hotel demand and revenue of 8 major tourist destinations in the US (e.g., Los Angeles, Orlando ...).
2 papers · 0 benchmarks
HuTics (Human Deictic Gestures Dataset)
HuTics contains 2040 images showing how humans use deictic gestures to interact with various daily-life objects.
2 papers · 0 benchmarks
Timely and effective response to humanitarian crises requires quick and accurate analysis of large amounts of text data, a process that can highly benefit from expert-assisted NLP systems trained on validated and annotated data in the…
2 papers · 0 benchmarks
Human Simulacra is a virtual character dataset that contains 129k texts across 11 virtual characters, with each character having unique attributes, biographies, and stories.
2 papers · 0 benchmarks
I2-2000FPS is the first high-speed video dataset offering an unprecedented temporal resolution of 2000 frames per second (fps).
2 papers · 0 benchmarks
IACC.3 (Internet Archive videos (IACC.3) under Creative Commons licenses.)
The IACC.3 dataset is approximately 4600 Internet Archive videos (144 GB, 600 h) with Creative Commons licenses in MPEG-4/H.264 format with duration ranging from 6.5 min to 9.5 min and a mean duration of almost 7.8 min.
2 papers · 0 benchmarks
IBL-NeRF Dataset.
2 papers · 0 benchmarks
IHDS (Indian Human Developement Survey)
IHDS is a nationally representative, multi-topic panel survey of 41,554 households in 1503 villages and 971 urban neighborhoods across India.
2 papers · 0 benchmarks
Bearing acceleration data from three run-to-failure experiments on a loaded shaft.
2 papers · 0 benchmarks
The IRFL dataset consists of idioms, similes, and metaphors with matching figurative and literal images, as well as two novel tasks of multimodal figurative understanding and preference.
2 papers · 2 benchmarks
The ISOT Fake News dataset is a compilation of several thousands fake news and truthful articles, obtained from different legitimate news sites and sites flagged as unreliable by Politifact.com.
2 papers · 0 benchmarks
Im4Sketch is a large-scale dataset with shape-oriented set of classes for image-to-sketch generalization .
2 papers · 1 benchmark
InVar-100 (Industrial Objects in Varied Contexts)
The Industrial Objects in Varied Contexts (InVar) Dataset was internally produced by our team and contains 100 objects in 20800 total images (208 images per class).
2 papers · 0 benchmarks
Inception Computational Graph (CG) family introduced in "GENNAPE: Towards Generalized Neural Architecture Performance Estimators", accepted to AAAI-23.
2 papers · 0 benchmarks
IndiaPoliceEvents is a corpus of 21,391 sentences from 1,257 English-language Times of India articles about events in the state of Gujarat during March 2002.
2 papers · 0 benchmarks
Industry Biscuit (Cookie) dataset (Industrial style dataset for the anomaly detection)
The Industrial Biscuits (Cookie) dataset is our internal dataset designed for the anomaly detection task, which captures Tarallini biscuits.
2 papers · 0 benchmarks
InfiniteBench (∞Bench: Extending Long Context Evaluation Beyond 100K Tokens)
Introduction Welcome to InfiniteBench, a cutting-edge benchmark tailored for evaluating the capabilities of language models to process, understand, and reason over super long contexts (100k+ tokens).
2 papers · 0 benchmarks
The Insider Threat Test Dataset is a collection of synthetic insider threat test datasets that provide both background and malicious actor synthetic data.
2 papers · 1 benchmark
InspiRe (Inspiring and non-inspiring posts from Reddit)
We analyze social media posts to tease out what makes a post inspiring and what topics are inspiring.
2 papers · 0 benchmarks
Instantiation is a dataset for the task of instantiation detection
2 papers · 0 benchmarks
IoT Traffic Traces (IoT Traffic Traces - Data Collected for IEEE TMC 2018)
IOT TRAFFIC TRACES Data Collected for IEEE TMC 2018 Cite our data A.
2 papers · 0 benchmarks
Jam-ALT (JamALT: A Formatting-Aware Lyrics Transcription Benchmark)
JamALT is a revision of the JamendoLyrics dataset (80 songs in 4 languages), adapted for use as an automatic lyrics transcription (ALT) benchmark.
2 papers · 5 benchmarks
KGRC-RDF-star is an RDF-star dataset converted from KGRC-RDF, which is a Knowledge graph dataset of novel stories.
2 papers · 0 benchmarks
We introduce KPI-EDGAR, a novel dataset for Joint Named Entity Recognition and Relation Extraction building on financial reports uploaded to the Electronic Data Gathering, Analysis, and Retrieval (EDGAR) system, where the main objective is…
2 papers · 1 benchmark
KanHope (Kannada Hope speech dataset)
KanHope is a code mixed hope speech dataset for equality, diversity, and inclusion in Kannada, an under-resourced Dravidian language.
2 papers · 1 benchmark
Kvasir-Capsule dataset is the largest publicly released VCE dataset.
2 papers · 0 benchmarks
This dataset is based on the LFM-1b [ and the Cultural LFM-1b [2] datasets.
2 papers · 0 benchmarks
LOFAR RFI Detection (Low-Frequency Array (LOFAR) Radio Frequency Interference Detection)
This dataset contains simulated and expert-labelled spectrograms from two radio telescopes: the Hydrogen Epoch of Reionization Array (HERA) in South Africa and the Low-Frequency Array (LOFAR) in the Netherlands.
2 papers · 1 benchmark
LUMA (Learning from Uncertain and Multimodal Data)
LUMA is a multimodal dataset that consists of audio, image, and text modalities.
2 papers · 0 benchmarks
Dataset of validated OCT and Chest X-Ray images described and analyzed in "Deep learning-based classification and referral of treatable human diseases".
2 papers · 0 benchmarks
The Large-Scale CLIR Dataset is a retrieval dataset built for Cross-Language Information Retrieval (CLIR).
2 papers · 0 benchmarks
This dataset presents a set of large-scale ridesharing Dial-a-Ride Problem (DARP) instances.
2 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.