Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 54 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 2545–2592 of 3,998
EGO-CH-Gaze (Learning to Detect Attended Objects in Cultural Sites with Gaze Signals and Weak Object Supervision)
To study the problem of weakly supervised attended object detection in cultural sites, we collected and labeled a dataset of egocentric images acquired from subjects visiting a cultural site.
1 paper · 0 benchmarks
EHT (The English Headline Treebank)
The English Headline Treebank (EHT) is an English headline treebank of 1,055 manually annotated and adjudicated universal dependency (UD) syntactic dependency trees to encourage research in improving NLP pipelines for English headlines.
1 paper · 0 benchmarks
For Emotion Interpretation task
1 paper · 2 benchmarks
ELITR Minuting Corpus in JSON format.
1 paper · 0 benchmarks
ELMTEX Dataset (ELMTEX Dataset: Fine-Tuning Large Language Models for Structured Clinical Information Extraction)
We introduced a new dataset of clinical report summaries, annotated with structured information across 15 categories.
1 paper · 0 benchmarks
EMMA (An Enhanced MultiModal ReAsoning Benchmark)
We introduce EMMA (Enhanced MultiModal reAsoning), a benchmark targeting organic multimodal reasoning across mathematics, physics, chemistry, and coding.
1 paper · 0 benchmarks
first everyday task dataset featuring COT outputs, diverse task designs, detailed re-plan processes, along with SFT and DPO sub-datasets.
1 paper · 0 benchmarks
ERD (Educational Resource Discovery)
ERD (Educational Resource Discovery) is a corpus of 39,728 manually labeled web resources and 659 queries from NLP, Computer Vision (CV), and Statistics (STATS) for educational resource discovery.
1 paper · 0 benchmarks
ESA-AD (European Space Agency Dataset for Anomaly Detection in Satellite Telemetry)
ESA Anomaly Dataset is the first large-scale, real-life satellite telemetry dataset with curated anomaly annotations originated from three ESA missions.
1 paper · 0 benchmarks
Dataset Card for ESG/DLT Named Entity Recognition Dataset This dataset contains named entities related to Distributed Ledger Technology (DLT) and Environmental, Social, and Governance (ESG) topics created to support research in these areas…
1 paper · 0 benchmarks
We present ESG-FTSE, the first corpus comprised of news articles with Environmental, Social and Governance (ESG) relevance annotations.
1 paper · 0 benchmarks
ESP (Evaluation for Styled Prompt)
ESP dataset (Evaluation for Styled Prompt dataset) is a benchmark for zero-shot domain-conditional caption generation.
1 paper · 0 benchmarks
Provide: A multimodal ESSVP dataset is built with 224×224 size RGB images and 59-channel EEG data.
1 paper · 0 benchmarks
EUCA dataset description Associated Paper: EUCA: the End-User-Centered Explainable AI Framework Authors: Weina Jin, Jianyu Fan, Diane Gromala, Philippe Pasquier, Ghassan Hamarneh Introduction: EUCA dataset is for modelling personalized or…
1 paper · 0 benchmarks
EUEN17037 Daylight and View Standard Test Dataset.
1 paper · 0 benchmarks
Dataset Description EUROPA is a dataset designed for training and evaluating multilingual keyphrase generation models in the legal domain.
1 paper · 0 benchmarks
The EVI dataset is a challenging, multilingual spoken-dialogue dataset with 5,506 dialogues in English, Polish, and French.
1 paper · 3 benchmarks
This is an assembly dataset built on top of ShellcodeIA32, a dataset for automatically generating assembly from natural language descriptions that consists of 3,200 assembly instructions, commented in the English language, which were…
1 paper · 0 benchmarks
This dataset contains samples to generate Python code for security exploits.
1 paper · 0 benchmarks
A Zero-Shot Sketch-based Inter-Modal Object Retrieval Scheme for Remote Sensing Images WITH the advancement in sensor technology, huge amounts of data are being collected from various satellites.
1 paper · 0 benchmarks
The dataset, generated from a scientific simulation, consists of a time series (251 steps) of 3D scalar fields on a spherical 180x201x360 grid covering 500 Myr of geological time.
1 paper · 0 benchmarks
A large dataset of over 18,000,000 English tweets posted by ∼7K echo users was constructed in the following manner: 1.
1 paper · 0 benchmarks
Echo Corpus (Arviv et al, 2021) infused with information from KnowledJe (Halevy, 2023).
1 paper · 0 benchmarks
Educational Grade School Math (EGSM) contains 2,093 question/answer pairs generated by MATHWELL, a reference-free educational grade school math word problem generator that outputs a word problem and Program of Thought (PoT) solution based…
1 paper · 0 benchmarks
EgoMon (Egomon Gaze & Video dataset)
EgoMon Gaze & Video Dataset is an Egocentric (first person) Dataset that consists of 7 videos of 30 minutes, more or less, each one of them.
1 paper · 0 benchmarks
This is a detailed description of the dataset, a data sheet for the dataset as proposed by Gebru et al.
1 paper · 0 benchmarks
EmoFilm (Emotional speech from Films)
EmoFilm is a multilingual emotional speech corpus comprising 1115 audio instances produced in English, Italian, and Spanish languages.
1 paper · 0 benchmarks
EmoSpeech contains keywords with diverse emotions and background sounds, presented to explore new challenges in audio analysis.
1 paper · 0 benchmarks
Predictions of energy consumption are crucial for energy retailers to minimize deviations from energy acquired in the day-ahead market and the actual consumption of their customers.
1 paper · 0 benchmarks
Benchmark to evaluate the capability of LMs to consolidate and recall information from multiple training documents.
1 paper · 0 benchmarks
This repository contains three graph datasets for the UE traffic assignment problem on Sioux-Falls, Eastern-Massachusetts and Anaheim networks in both dgl and pyg formats.
1 paper · 0 benchmarks
Provide: a high-level explanation of the dataset characteristics explain motivations and summary of its content potential use cases of the dataset Collection of Error Grid data files.
1 paper · 0 benchmarks
Essays (Stream-of-consciousness Essays)
J.
1 paper · 1 benchmark
EuroSAT-C is an open-source data set comprising algorithmically generated corruptions applied to the EuroSAT test set following the concept of ImageNet-C.
1 paper · 0 benchmarks
The ConcoDisco Corpus is an English-French parallel corpus with discourse relations (DRs) and discourse connectives (DCs) annotations.
1 paper · 0 benchmarks
A corpus designed in analogy to the well-established English ISEAR emotion dataset.
1 paper · 0 benchmarks
EventEA is an event-centric entity alignment dataset, harvested from EventKG, DBpedia and Wikidata.
1 paper · 0 benchmarks
Intermediate annotations from the FEVER dataset that describe original facts extracted from Wikipedia and the mutations that were applied, yielding the claims in FEVER.
1 paper · 0 benchmarks
ExBAN (ExBAN Corpus (Explanations for BAyesian Networks))
The ExBAN dataset: a corpus of NL explanations generated by crowd-sourced participants presented with the task of explaining simple Bayesian Network (BN) graphical representations.
1 paper · 0 benchmarks
ExPUNations is a humor dataset with such extensive and fine-grained annotations specifically for puns.
1 paper · 0 benchmarks
This is an example data set for a hypothetical electronic products supply network.
1 paper · 0 benchmarks
This repository presents the dataset used in the PerfCam's original paper.
1 paper · 0 benchmarks
This dataset is being used to evaluate PerfSim accuracy and speed against a real deployment in a Kubernetes cluster based on sfc-stress workloads.
1 paper · 0 benchmarks
Expository-Prose-V1 is a collection of specially-curated corpora gathered from diverse sources, ranging from research papers (arXiv) to European Parliament proceedings (EuroParl).
1 paper · 0 benchmarks
Neural network model files and Madgraph event generator outputs used as inputs to the results presented in the paper "Learning to discover: expressive Gaussian mixture models for multi-dimensional simulation and parameter inference in the…
1 paper · 0 benchmarks
A dataset of abdominal CT studies in NifTi format from the open-source medical data repository Medical Decathlon was utilized.
1 paper · 1 benchmark
The proposed Extended-YouTube Faces (E-YTF) is an extension of the famous YouTube Faces (YTF) dataset and is specifically designed to further push the challenges of face recognition by addressing the problem of open-set face identification…
1 paper · 0 benchmarks
The dataset X of this work is an extension of the heartSeg dataset.
1 paper · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.