Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 29 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 1345–1392 of 3,998

DSTC7 Task 2 (Dialog System Technology Challenges Task 2)
DSTC Task 2 is a dataset and task for end-to-end conversation modeling.
4 papers · 0 benchmarks
A large-scale anime image database with 4.2m+ images annotated with 130m+ text tags describing image contents in detail; it can be useful for machine learning purposes such as image recognition and generation.
4 papers · 0 benchmarks
Databricks Dolly 15k (databricks-dolly-15k)
Databricks Dolly 15k is a dataset containing 15,000 high-quality human-generated prompt / response pairs specifically designed for instruction tuning large language models.
4 papers · 0 benchmarks
The data set covers recordings of ripening fruit with labels of destructive measurements (fruit flesh firmness, sugar content and overall ripeness).
4 papers · 1 benchmark
The dataset consists of over 350,000 public domain patent drawings collected from the United States Patent and Trademark Office (USPTO).
4 papers · 1 benchmark
DEVAI is a benchmark of 55 realistic AI development tasks.
4 papers · 0 benchmarks
DiS-ReX is a multilingual dataset for distantly supervised (DS) relation extraction (RE).
4 papers · 0 benchmarks
DiaASQ (Conversational Aspect-based Sentiment Quadruple Extraction)
DiaASQ is a fine-grained Aspect-based Sentiment Analysis (ABSA) benchmark under the conversation scenario.
4 papers · 2 benchmarks
A new English-French test set for the evaluation of Machine Translation (MT) for informal, written bilingual dialogue.
4 papers · 1 benchmark
DiaMOS Plant (A Dataset for Diagnosis and Monitoring Plant Disease)
Abstract The classification and recognition of foliar diseases is an increasingly developing field of research, where the concepts of machine and deep learning are used to support agricultural stakeholders.
4 papers · 0 benchmarks
ECG-Image-Database (Digitization and Classification of ECG Images: The George B. Moody PhysioNet Challenge 2024)
The George B.
4 papers · 1 benchmark
EDGE-IIOTSET (A NEW COMPREHENSIVE REALISTIC CYBER SECURITY DATASET OF IOT AND IIOT APPLICATIONS: CENTRALIZED AND FEDERATED LEARNING)
ABSTRACT In this project, we propose a new comprehensive realistic cyber security dataset of IoT and IIoT applications, called Edge-IIoTset, which can be used by machine learning-based intrusion detection systems in two different modes,…
4 papers · 0 benchmarks
ETHEC (ETH Entomological Collection (ETHEC) Dataset)
It includes 47,978 butterfly images with a 4-level label-hierarchy.
4 papers · 0 benchmarks
EVA (EVA-7K dataset)
The dataset contains 7000 videos: native, altered and exchanged through social platforms.
4 papers · 0 benchmarks
EdAcc (Edinburgh International Accents of English Corpus)
The Edinburgh International Accents of English Corpus (EdAcc) is a new automatic speech recognition (ASR) dataset composed of 40 hours of English dyadic conversations between speakers with a diverse set of accents.
4 papers · 0 benchmarks
EurekaAlert (Eureka Alert)
This dataset contains around 5000 scholarly articles and their corresponding easy summary from eureka alert blog, the dataset can be used for the combined task of summarization and simplification.
4 papers · 2 benchmarks
FACTIFY (a dataset on multi-modal fact verification)
FACTIFY is a dataset on multi-modal fact verification.
4 papers · 0 benchmarks
FR-FS (Fall Recognition in Figure Skating)
The FR-FS dataset contains 417 videos collected from FIV dataset and Pingchang 2018 Winter Olympic Games.
4 papers · 0 benchmarks
FRMT (Few-shot Region-aware Machine Translation)
FRMT is a dataset and evaluation benchmark for Few-shot Region-aware Machine Translation, a type of style-targeted translation.
4 papers · 4 benchmarks
FacetSum is a faceted summarization dataset for scientific documents.
4 papers · 1 benchmark
The FakeMusicCaps dataset contains total of 27605 10 seconds music tracks corresponding to almost 77 hours, generated using 5 different Text-To-Music (TTM) models.
4 papers · 0 benchmarks
FewSOL (A Dataset for Few-Shot Object Learning in Robotic Environments)
The Few-Shot Object Learning (FewSOL) dataset can be used for object recognition with a few images per object.
4 papers · 0 benchmarks
FiNER-139 is comprised of 1.1M sentences annotated with eXtensive Business Reporting Language (XBRL) tags extracted from annual and quarterly reports of publicly-traded companies in the US.
4 papers · 0 benchmarks
The first NER dataset in the field of traffic, which is to extract the characteristics and attributes of the vehicle on the road.
4 papers · 2 benchmarks
A dataset of high resolution, textured scans of articulated left feet, useful for 3D shape representation learning.
4 papers · 0 benchmarks
Fruits 360 (A dataset of images containing fruits, vegetables, nuts and seeds)
Fruits-360 dataset: A dataset of images containing fruits, vegetables, nuts and seeds Version: 2025.03.24.0 Content The following fruits, vegetables and nuts and are included: Apples (different varieties: Crimson Snow, Golden, Golden-Red,…
4 papers · 0 benchmarks
GDSC (Genomics of Drug Sensitivity in Cancer)
We have characterized 1000 human cancer cell lines and screened them with 100s of compounds.
4 papers · 1 benchmark
GeoCoV19 is a large-scale Twitter dataset containing more than 524 million multilingual tweets.
4 papers · 0 benchmarks
Goal is a novel dataset of football (or 'soccer') highlights videos with transcribed live commentaries in English.
4 papers · 0 benchmarks
HOMER (Household Object Movements from Everyday Routines)
The Household Object Movements from Everyday Routines (HOMER) dataset is composed of routine behaviors for five households, spanning 50 days for the train split and 10 days for test split.
4 papers · 0 benchmarks
Hate speech has become one of the most significant issues in modern society, with implications in both the online and offline worlds.
4 papers · 1 benchmark
Images with paired ground-truth caption hierarchies
4 papers · 0 benchmarks
This dataset contains five notable histological artifacts: blur, blood (hemorrhage), air bubbles, folded tissue, and damaged tissue.
4 papers · 1 benchmark
Human-Animal-Cartoon (HAC) dataset consists of seven actions (‘sleeping’, ‘watching tv’, ‘eating’, ‘drinking’, ‘swimming’, ‘running’, and ‘opening door’) performed by humans, animals, and cartoon figures, forming three different domains.
4 papers · 0 benchmarks
The IS-A dataset is a dataset of relations extracted from a medical ontology.
4 papers · 0 benchmarks
ImgEdit is a large-scale, high-quality image-editing dataset comprising 1.2 million carefully curated edit pairs, which contain both novel and complex single-turn edits, as well as challenging multi-turn tasks.
4 papers · 0 benchmarks
InferWiki is a Knowledge Graph Completion (KGC) dataset that improves upon existing benchmarks in inferential ability, assumptions, and patterns.
4 papers · 0 benchmarks
You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior.
4 papers · 1 benchmark
JobStack is a new corpus for de-identification of personal data in job vacancies on Stackoverflow.
4 papers · 0 benchmarks
K-Lane (KAIST-Lane)
KAIST-Lane (K-Lane) is the world’s first and the largest public urban road and highway lane dataset for Lidar.
4 papers · 1 benchmark
Human Activity Recognition (HAR) refers to the capacity of machines to perceive human actions.
4 papers · 0 benchmarks
Kitsune Network Attack Dataset This is a collection of nine network attack datasets captured from a either an IP-based commercial surveillance system or a network full of IoT devices.
4 papers · 0 benchmarks
This is the dataset for knowledge editing.
4 papers · 0 benchmarks
KodCode-V1 (KodCode/KodCode-V1)
KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks.
4 papers · 0 benchmarks
Kosp2e (read as kospi'), is a corpus that allows Korean speech to be translated into English text in an end-to-end manner
4 papers · 0 benchmarks
The LIAR dataset has been widely followed by fake news detection researchers since its release, and along with a great deal of research, the community has provided a variety of feedback on the dataset to improve it.
4 papers · 1 benchmark
LIMUC (Labeled Images for Ulcerative Colitis)
The LIMUC dataset is the largest publicly available labeled ulcerative colitis dataset that compromises 11276 images from 564 patients and 1043 colonoscopy procedures.
4 papers · 1 benchmark
Laptop-ACOS is a brand new Laptop dataset collected from the Amazon platform in the years 2017 and 2018 (covering ten types of laptops under six brands such as ASUS, Acer, Samsung, Lenovo, MBP, MSI, and so on).
4 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.