Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 6 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 241–288 of 3,998
WikiLarge comprise 359 test sentences, 2000 development sentences and 300k training sentences.
69 papers · 0 benchmarks
BeaverTails is a dataset aimed at fostering research on safety alignment in large language models (LLMs).
68 papers · 0 benchmarks
LIVE-VQC (LIVE Video Quality Challenge (VQC) Database)
The great variations of videographic skills in videography, camera designs, compression and processing protocols, communication and bandwidth environments, and displays leads to an enormous variety of video impairments.
68 papers · 1 benchmark
The PlantVillage dataset consists of 54303 healthy and unhealthy leaf images divided into 38 categories by species and disease.
68 papers · 1 benchmark
HatEval (SemEval 2019 Task 5: Multilingual Detection of Hate Speech Against Immigrants and Women in Twitter)
Hate Speech is commonly defined as any communication that disparages a person or a group on the basis of some characteristic such as race, color, ethnicity, gender, sexual orientation, nationality, religion, or other characteristics.
67 papers · 1 benchmark
CoVoST2 (Common Voice Speech-To-Text 2)
End-to-end speech-to-text translation (ST) has recently witnessed an increased interest given its system simplicity, lower inference latency and less compounding errors compared to cascaded ST (i.e.
66 papers · 0 benchmarks
This project contains natural language data for human-robot interaction in home domain which we collected and annotated for evaluating NLU Services/platforms.
66 papers · 3 benchmarks
NAB (Numenta Anomaly Benchmark)
The First Temporal Benchmark Designed to Evaluate Real-time Anomaly Detectors Benchmark The growth of the Internet of Things has created an abundance of streaming data.
66 papers · 1 benchmark
Text corpus with almost one billion words of training data for statistical language modeling benchmarking.
66 papers · 0 benchmarks
WikiHop is a multi-hop question-answering dataset.
66 papers · 2 benchmarks
ACE 2005 (ACE 2005 Multilingual Training Corpus)
ACE 2005 Multilingual Training Corpus contains the complete set of English, Arabic and Chinese training data for the 2005 Automatic Content Extraction (ACE) technology evaluation.
65 papers · 8 benchmarks
DBP15k contains four language-specific KGs that are respectively extracted from English (En), Chinese (Zh), French (Fr) and Japanese (Ja) DBpedia, each of which contains around 65k-106k entities.
65 papers · 3 benchmarks
This dataset consists of (human-written) NBA basketball game summaries aligned with their corresponding box- and line-scores.
65 papers · 5 benchmarks
AIDA CoNLL-YAGO contains assignments of entities to the mentions of named entities annotated for the original CoNLL 2003 entity recognition task.
64 papers · 0 benchmarks
MagicBrush is a manually-annotated instruction-guided image editing dataset covering diverse scenarios single-turn, multi-turn, mask-provided, and mask-free editing.
64 papers · 0 benchmarks
OSCAR or Open Super-large Crawled ALMAnaCH coRpus is a huge multilingual corpus obtained by language classification and filtering of the Common Crawl corpus using the goclassy architecture.
64 papers · 0 benchmarks
The TED-LIUM corpus consists of English-language TED talks.
64 papers · 2 benchmarks
XL-Sum is a comprehensive and diverse dataset for abstractive summarization comprising 1 million professionally annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics.
64 papers · 0 benchmarks
ESD (Emotional Speech Database)
ESD is an Emotional Speech Database for voice conversion research.
63 papers · 0 benchmarks
We propose Localized Narratives, a new form of multimodal image annotations connecting vision and language.
63 papers · 3 benchmarks
SLAKE is an English-Chinese bilingual dataset consisting of 642 images and 14,028 question-answer pairs for training and testing Med-VQA systems.
63 papers · 0 benchmarks
Sentence Compression is a dataset where the syntactic trees of the compressions are subtrees of their uncompressed counterparts, and hence where supervised systems which require a structural alignment between the input and output can be…
63 papers · 0 benchmarks
ComplexWebQuestions is a dataset for answering complex questions that require reasoning over multiple web snippets.
62 papers · 2 benchmarks
Dataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary.
62 papers · 0 benchmarks
CIRR (Compose Image Retrieval on Real-life images)
Composed Image Retrieval (or, Image Retreival conditioned on Language Feedback) is a relatively new retrieval task, where an input query consists of an image and short textual description of how to modify the image.
61 papers · 3 benchmarks
The LIP (Look into Person) dataset is a large-scale dataset focusing on semantic understanding of a person.
61 papers · 1 benchmark
Node classification on Penn94
60 papers · 2 benchmarks
ToTTo is an open-domain English table-to-text dataset with over 120,000 training examples that proposes a controlled generation task: given a Wikipedia table and a set of highlighted table cells, produce a one-sentence description.
60 papers · 1 benchmark
Subset and preprocessed version of Chemical reactions from US patents (1976-Sep2016) by Daniel Lowe.
60 papers · 1 benchmark
CELEX database comprises three different searchable lexical databases, Dutch, English and German.
59 papers · 0 benchmarks
WOS (Web of Science Dataset)
Web of Science (WOS) is a document classification dataset that contains 46,985 documents with 134 categories which include 7 parents categories.
59 papers · 4 benchmarks
CANARD (A Dataset for Question-in-Context Rewriting)
CANARD is a dataset for question-in-context rewriting that consists of questions each given in a dialog context together with a context-independent rewriting of the question.
58 papers · 1 benchmark
CCNet is a dataset extracted from Common Crawl with a different filtering process than for OSCAR.
58 papers · 0 benchmarks
EmoryNLP comprises 97 episodes, 897 scenes, and 12,606 utterances, where each utterance is annotated with one of the seven emotions borrowed from the six primary emotions in the Willcox (1982)’s feeling wheel, sad, mad, scared, powerful,…
58 papers · 1 benchmark
We introduce a dataset of 147 object categories containing over 6000 images that are suitable for the few-shot counting task.
58 papers · 4 benchmarks
The Implicit Hate corpus is a dataset for hate speech detection with fine-grained labels for each message and its implication.
58 papers · 0 benchmarks
SQA3D (Situated Question Answering in 3D Scenes)
SQA3D is a dataset for embodied scene understanding, where an agent needs to comprehend the scene it situates from an first person's perspective and answer questions.
58 papers · 3 benchmarks
XD-Violence is a large-scale audio-visual dataset for violence detection in videos.
58 papers · 2 benchmarks
Assembly101 is a new procedural activity dataset featuring 4321 videos of people assembling and disassembling 101 "take-apart" toy vehicles.
57 papers · 4 benchmarks
Europarl-ST is a multilingual Spoken Language Translation corpus containing paired audio-text samples for SLT from and into 9 European languages, for a total of 72 different translation directions.
57 papers · 0 benchmarks
Fluent Speech Commands is an open source audio dataset for spoken language understanding (SLU) experiments.
57 papers · 1 benchmark
MSLS (Mapillary Street-level Sequences Dataset)
The largest and most diverse dataset for lifelong place recognition from image sequences in urban and suburban settings.
57 papers · 1 benchmark
WHAMR! (WHAM! with synthetic reverberated sources)
WHAMR!
57 papers · 3 benchmarks
Data Set Information: Extraction was done by Barry Becker from the 1994 Census database.
56 papers · 2 benchmarks
BEAT (Body-Expression-Audio-Text)
BEAT has i) 76 hours, high-quality, multi-modal data captured from 30 speakers talking with eight different emotions and in four different languages, ii) 32 millions frame-level emotion and semantic relevance annotations.
56 papers · 1 benchmark
PA-100K is a recent-proposed large pedestrian attribute dataset, with 100,000 images in total collected from outdoor surveillance cameras.
56 papers · 1 benchmark
The Re-TACRED dataset is a significantly improved version of the TACRED dataset for relation extraction.
56 papers · 1 benchmark
MuTual is a retrieval-based dataset for multi-turn dialogue reasoning, which is modified from Chinese high school English listening comprehension test data.
55 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.