Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 75 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 3553–3600 of 3,998

We present the Webis-STEREO-21 dataset, a massive collection of Scientific Text Reuse in Open-access publications.
1 paper · 0 benchmarks
This corpus contains preprocessed posts from the Reddit dataset, suitable for abstractive summarization using deep learning.
1 paper · 0 benchmarks
Well-being Dataset (Cambridge Well-being Dataset for Psychological Distress Analysis)
The dataset is a private dataset collected for automatic analysis of psychological distress.
1 paper · 1 benchmark
Werewolf Among Us is a dataset multimodal dataset for modeling persuasion behaviors.
1 paper · 0 benchmarks
Manually labelled dataset of bird recordings from the species of interest inhabiting in the wetlands of the "Aiguamolls del Empord\{a}" natural park in Girona, Spain.
1 paper · 0 benchmarks
WhenAct (Temporal Human Action Localization in Lifestyle Vlogs)
We consider the task of temporal human action localization in lifestyle vlogs.
1 paper · 0 benchmarks
WiFiCam dataset for through-wall imaging based on WiFi channel state information.
1 paper · 0 benchmarks
WiRLD (Wikidata Reference Logo Dataset)
The Wikidata Reference Logo Dataset (WiRLD), a comprehensive collection of reference logos specifically designed to address the challenges of large-scale logo identification.
1 paper · 0 benchmarks
WiRLD_ (Wikidata Reference Logo Dataset)
The Wikidata Reference Logo Dataset (WiRLD), a comprehensive collection of reference logos specifically designed to address the challenges of large-scale logo identification.
1 paper · 0 benchmarks
WiTA (Writing in The Air)
WiTA (Writing in The Air) is a dataset for the challenging writing in the air (WiTA) task -- an elaborate task bridging vision and NLP.
1 paper · 0 benchmarks
The Wiki-Flick Event dataset for cross-modal event retrieval is a well-labelled but weakly-aligned dataset collected for cross-modality event retrieval.
1 paper · 0 benchmarks
Wiki-Reliability is the first dataset of English Wikipedia articles annotated with a wide set of content reliability issues.
1 paper · 0 benchmarks
Wiki-en is an annotated English dataset for domain detection extracted from Wikipedia.
1 paper · 0 benchmarks
WikiBanEvasion (Wikipedia Ban Evasion Dataset)
A dataset comprising 8,551 ban evasion pairs on Wikipedia, where each pair comprises a parent account and the child account.
1 paper · 0 benchmarks
WikiBioCTE is a dataset for controllable text edition based on the existing dataset WikiBio (originally created for table-to-text generation).
1 paper · 0 benchmarks
WikiChurches is a dataset for architectural style classification, consisting of 9,485 images of church buildings.
1 paper · 0 benchmarks
WikiDes is a dataset for generating descriptions of Wikidata from Wikipedia paragraphs.
1 paper · 0 benchmarks
WikiEvalFacts This is an augmented version of the WikiEval dataset which additionally includes generated fact statements for each QA pair and human annotation of their truthfulness against each answer type.
1 paper · 0 benchmarks
WikiFactDiff is a dataset designed as a resource to perform atomic factual knowledge updates on language models, with the goal of aligning them with current knowledge.
1 paper · 0 benchmarks
WikiMulti (WikiMulti: a Corpus for Cross-Lingual Summarization)
wikimulti is a dataset for cross-lingual summarization based on Wikipedia articles in 15 languages.
1 paper · 0 benchmarks
WikiOFGraph (Wikipedia Ontology-Free Graph-Text)
a high-level explanation of the dataset characteristics We introduce WikiOFGraph, a novel large-scale, domain-diverse dataset synthesized by LLMs, ensuring superior graph-text consistency to advance general-domain graph-to-text generation.
1 paper · 1 benchmark
WikiPII, an automatically labeled dataset composed of Wikipedia biography pages, annotated for personal information extraction.
1 paper · 0 benchmarks
WikiWeb2M (Wikipedia Webpage 2M)
Wikipedia Webpage 2M (WikiWeb2M) is a multimodal open source dataset consisting of over 2 million English Wikipedia articles.
1 paper · 0 benchmarks
Wikipedia is the largest and most read online free encyclopedia currently existing.
1 paper · 0 benchmarks
Wikipedia users activity for two language editions, Portuguese and Italian, for up to 8 January 2020.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
WinoPron is a novel dataset of Winogender-like template pairs in English, which fixes inconsistencies in Winogender Schemas and contains balanced template pairs for pronoun forms in 3 grammatical cases, which we find impacts performance…
1 paper · 0 benchmarks
WordNet-feelings, is an affective dataset that identifies 3664 word senses as feelings, and associates each of these with one of the 9 categories of feeling.
1 paper · 0 benchmarks
We present the World Wide Dishes dataset which seeks to assess disparities in representations of food through a decentralised data collection effort to gather perspectives directly from people with a wide variety of backgrounds from around…
1 paper · 0 benchmarks
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
We provide a new data set XWikiRef for the task of Cross-lingual Multi-document Summarization.
1 paper · 0 benchmarks
YADL (Yet Another Data LAke)
Files composing the YADL data lake, for the paper "Retrieve, Merge, Predict: Augmenting Tables with Data Lakes (Experiment, Analysis & Benchmark Paper)" Archives provided here follow the notation used for the experiments, which is…
1 paper · 0 benchmarks
YIM Dataset (Yeast Cells in Microstructures Dataset)
An instance segmentation dataset of yeast cells in microstructures.
1 paper · 0 benchmarks
YJMob100K (YJMob100K: City-Scale and Longitudinal Dataset of Anonymized Human Mobility Trajectories)
Modeling and predicting human mobility trajectories in urban areas is an essential task for various applications including transportation modeling, disaster management, and urban planning.
1 paper · 0 benchmarks
Yeast colony morphologies (Quantifying yeast colony morphologies with feature engineering from time-lapse photography)
Data for the paper entitled Quantifying yeast colony morphologies with feature engineering from time-lapse photography by A.
1 paper · 0 benchmarks
YesBut Dataset (https://yesbut-dataset.github.io) Understanding satire and humor is a challenging task for even current Vision-Language models.
1 paper · 0 benchmarks
YorkTag provides pairs of sharp/blurred images containing fiducial markers and is proposed to train and qualitatively and quantitatively evaluate our model.
1 paper · 0 benchmarks
YouwikiHow is a dataset for Weakly-Supervised temporal Article Grounding (WSAG).
1 paper · 0 benchmarks
ZeroKBC is comprehensive benchmark that covers all scenarios of zero-shot Knowledge Base Completion (KBC) task.
1 paper · 0 benchmarks
ai4st SLR (Research on AI for Software Testing Research, 2020-2025)
To check the validity of the ai4st ontology, an adapted, lightweight systematic literature review (SLR) was conducted to analyse related research.
1 paper · 0 benchmarks
alpha-matte MFIF dataset (alpha-matte multi-focus image fusion dataset)
A large-scale training dataset suffering from the defocus spread effect (DSE) is synthesized by applying an α-matte boundary defocus model to the VOC 2012 dataset.
1 paper · 0 benchmarks
approved_drug_target (Approved Drug SMILES and Protein Sequence Dataset)
This dataset provides a curated collection of approved drug Simplified Molecular Input Line Entry System (SMILES) strings and their associated protein sequences.
1 paper · 0 benchmarks
This is a second public release of the arXMLiv dataset generated by the KWARC research group.
1 paper · 0 benchmarks
arXiv Categories (arXiv Categories Multi-label Text Classification Dataset)
This is a dataset of scientific documents derived from arXiv.
1 paper · 0 benchmarks
bSDD (buildingSMART Data Dictionary)
The buildingSMART Data Dictionary (bSDD) is an online service that hosts classifications and their properties, allowed values, units and translations.
1 paper · 0 benchmarks
In this repository, we provide the set-up files and output files of 5 behavioral observation data entry applications.
1 paper · 0 benchmarks
bigscience/P3 (bigscience/P3, split='ai2_arc_ARC_Challenge_pick_the_most_correct_option')
This datasets consists of challenging reasoning questions in multiple choice format.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.