Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 20 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 913–960 of 3,998

Global Voices is a multilingual dataset for evaluating cross-lingual summarization methods.
9 papers · 0 benchmarks
Global Wheat (Global Wheat Head Dataset 2020)
Global WHEAT Dataset is the first large-scale dataset for wheat head detection from field optical images.
9 papers · 0 benchmarks
HQ-WMCA (High-Quality Wide Multi-Channel Attack database)
The High-Quality Wide Multi-Channel Attack database (HQ-WMCA) database consists of 2904 short multi-modal video recordings of both bona-fide and presentation attacks.
9 papers · 0 benchmarks
Paper | Github | Dataset| Model As a part of our research efforts toward making LLMs more safe for public use, we create HarmfulQA i.e.
9 papers · 1 benchmark
Home Action Genome is a large-scale multi-view video database of indoor daily activities.
9 papers · 2 benchmarks
Imagewoof is a subset of 10 dog breed classes from Imagenet.
9 papers · 0 benchmarks
Dataset for document shadow removal
9 papers · 0 benchmarks
KnowIT VQA is a video dataset with 24,282 human-generated question-answer pairs about The Big Bang Theory.
9 papers · 0 benchmarks
LDC2020T02 (Abstract Meaning Representation (AMR) Annotation Release 3.0)
Abstract Meaning Representation (AMR) Annotation Release 3.0 was developed by the Linguistic Data Consortium (LDC), SDL/Language Weaver, Inc., the University of Colorado's Computational Language and Educational Research group and the…
9 papers · 1 benchmark
MSP-Podcast (A large naturalistic speech emotional dataset)
The MSP-Podcast corpus contains speech segments from podcast recordings which are perceptually annotated using crowdsourcing.
9 papers · 4 benchmarks
MuSe-CaR (Multimodal Sentiment Analysis in Car Reviews)
The MuSe-CAR database is a large, multimodal (video, audio, and text) dataset which has been gathered in-the-wild with the intention of further understanding Multimodal Sentiment Analysis in-the-wild, e.g., the emotional engagement that…
9 papers · 0 benchmarks
NELA-GT-2018 is a dataset for the study of misinformation that consists of 713k articles collected between 02/2018-11/2018.
9 papers · 0 benchmarks
OntoGUM is an OntoNotes-like coreference dataset converted from GUM, an English corpus covering 12 genres using deterministic rules.
9 papers · 1 benchmark
PhenoBench (PhenoBench — A Large Dataset and Benchmarks for Semantic Image Interpretation in the Agricultural Domain)
The PhenoBench dataset contains multiple image segmentation challenges from the agricultural domain.
9 papers · 0 benchmarks
RadQA (A Question Answering Dataset to Improve Comprehension of Radiology Reports)
RadQA is a radiology question answering dataset with 3074 questions posed against radiology reports and annotated with their corresponding answer spans (resulting in a total of 6148 question-answer evidence pairs) by physicians.
9 papers · 1 benchmark
The SALMon dataset and benchmark was introduced in the paper "A Suite for Acoustic Language Model Evaluation", with the goal of evaluating the modelling abilities of speech language models with regards to different kinds of acoustic…
9 papers · 1 benchmark
Speech encompasses a wealth of information, including but not limited to content, paralinguistic, and environmental information.
9 papers · 0 benchmarks
SPARTQA (SPAtial Reasoning on Textual Question Answering)
SpartQA is a textual question answering benchmark for spatial reasoning on natural language text which contains more realistic spatial phenomena not covered by prior datasets and that is challenging for state-of-the-art language models…
9 papers · 0 benchmarks
SPARTQA - (SPAtial Reasoning on Textual Question Answering.)
We take advantage of the ground truth of NLVR images, design CFGs to generate stories, and use spatial reasoning rules to ask and answer spatial reasoning questions.
9 papers · 0 benchmarks
The SciTail dataset is an entailment dataset created from multiple-choice science exams and web sentences.
9 papers · 1 benchmark
TEMPO (Localizing Moments in Video with Temporal Language)
TEMPOral reasoning in video and language (TEMPO) is a dataset that consists of two parts: a dataset with real videos and template sentences (TEMPO - Template Language) which allows for controlled studies on temporal language, and a human…
9 papers · 0 benchmarks
UIIS (General Underwater Image Instance Segmentation dataset)
This is the first general Underwater Image Instance Segmentation (UIIS) dataset containing 4,628 images for 7 categories with pixel-level annotations for underwater instance segmentation task
9 papers · 1 benchmark
Leonardo Filipe Rodrigues Ribeiro, Pedro H.
9 papers · 1 benchmark
Have need seven multiple exposure ground truth images satisfying EV 0, ±1, ±2, ±3 for static scenes.
9 papers · 1 benchmark
VISUELLE is a repository build upon the data of a real fast fashion company, Nunalie, and is composed of 5577 new products and about 45M sales related to fashion seasons from 2016-2019.
9 papers · 1 benchmark
VLM2-Bench (VLM²-Bench)
VLM²-Bench: Benchmarking Vision-Language Models on Visual Cue Matching Description VLM²-Bench is the first comprehensive benchmark designed to evaluate vision-language models' (VLMs) ability to visually link matching cues across…
9 papers · 1 benchmark
We present a new large-scale human value dataset called ValueNet, which contains human attitudes on 21,374 text scenarios.
9 papers · 0 benchmarks
WiC-TSV (Words-in-Context: Target Sense Verification)
WiC-TSV is a new multi-domain evaluation benchmark for Word Sense Disambiguation.
9 papers · 2 benchmarks
The Zenseact Open Dataset (ZOD) is a large-scale and diverse multi-modal autonomous driving (AD) dataset, created by researchers at Zenseact.
9 papers · 0 benchmarks
The iWildCam2020-WILDS dataset is a variant of the iWildCam 2020 dataset.
9 papers · 1 benchmark
A scholarly data set with publications’ full-text, annotated in-text citations, and links to metadata.
9 papers · 0 benchmarks
3DCSR dataset (3D cross-source point cloud registration dataset)
Cross-source point cloud dataset for registration task.
8 papers · 0 benchmarks
ATLAS v2.0 (Anatomical Tracings of Lesions After Stroke Dataset version 2.0)
Accurate lesion segmentation is critical in stroke rehabilitation research for the quantification of lesion burden and accurate image processing.
8 papers · 1 benchmark
AmaSum is the largest abstractive opinion summarization dataset, consisting of more than 33,000 human-written summaries for Amazon products.
8 papers · 0 benchmarks
Amazon Toys & Games (Amazon Toys & Games 5-core)
This dataset includes reviews (ratings, text, helpfulness votes), product metadata (descriptions, category information, price, brand, and image features), and links (also viewed/also bought graphs).
8 papers · 1 benchmark
Amazon-Fraud (Multi-relational Graph Dataset for Amazon Fraudulent Account Detection)
Amazon-Fraud is a multi-relational graph dataset built upon the Amazon review dataset, which can be used in evaluating graph-based node classification, fraud detection, and anomaly detection models.
8 papers · 3 benchmarks
The Arena-Hard-Auto benchmark is an automatic evaluation tool for instruction-tuned Language Learning Models (LLMs)¹.
8 papers · 0 benchmarks
BB (Bacteria Biotope)
The Bacteria Biotope (BB) Task is part of the BioNLP Open Shared Tasks and meets the BioNLP-OST standards of quality, originality and data formats.
8 papers · 0 benchmarks
Brazil Air-Traffic
8 papers · 1 benchmark
CAD (Contextual Abuse Dataset)
Dataset of primarily English Reddit entries which addresses several limitations of prior work.
8 papers · 1 benchmark
Dataset [46 M] and readme: 42,306 movie plot summaries extracted from Wikipedia + aligned metadata extracted from Freebase, including: Movie box office revenue, genre, release date, runtime, and language Character names and aligned…
8 papers · 0 benchmarks
COCO-MIG (COCO-MIG benchmark)
The COCO-MIG benchmark (Common Objects in Context Multi-Instance Generation) is a benchmark used to evaluate the generation capability of generators on text containing multiple attributes of multi-instance objects.
8 papers · 1 benchmark
CoNLL-2000 is a dataset for dividing text into syntactically related non-overlapping groups of words, so-called text chunking.
8 papers · 0 benchmarks
A large commercial Ads Dataset includes 480K labeled query-ad pairwise data with structured information of image, title, seller, description, and so on.
8 papers · 1 benchmark
The DSSE-200 is a complex document layout dataset including various dataset styles.
8 papers · 0 benchmarks
DocNLI is a large-scale dataset for document-level NLI.
8 papers · 0 benchmarks
Duke Breast Cancer MRI (Dynamic contrast-enhanced magnetic resonance images of breast cancer patients with tumor locations)
Breast MRI scans of 922 cancer patients from Duke University, with tumor bounding box annotations, clinical, imaging, and many other features, and more.
8 papers · 0 benchmarks
EUR-Lex-Sum is a dataset for cross-lingual summarization.
8 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.