Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 24 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 1105–1152 of 3,998

This dataset contains around 10000 videos generated by various methods using the Prompt list.
6 papers · 1 benchmark
F-CelebA (10 tasks) (Federated-CelebA (10 tasks))
F-CelebA - This dataset is adapted from federated learning.
6 papers · 1 benchmark
FusedChat is an inter-mode dialogue dataset.
6 papers · 1 benchmark
This dataset is a new benchmark, grounded in real-world usages is developed to support more authentic and comprehensive evaluation of image editing models.
6 papers · 2 benchmarks
GTSinger (GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks)
The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited diversity of languages…
6 papers · 0 benchmarks
HRS-Bench (Holistic, Reliable, and Scalable Benchmark)
HRS-Bench is a concrete evaluation benchmark for T2I models that is Holistic, Reliable, and Scalable.
6 papers · 0 benchmarks
We introduce HumanEval-XL, a massively multilingual code generation benchmark specifically crafted to address this deficiency.
6 papers · 0 benchmarks
HurricaneEmo is an emotion dataset that contains 15,000 English tweets spanning three hurricanes: Harvey, Irma, and Maria.
6 papers · 0 benchmarks
IBM-Rank-30k (IBM-ArgQ-Rank-30kArgs)
The IBM-Rank-30k is a dataset for the task of argument quality ranking.
6 papers · 0 benchmarks
JParaCrawl is a parallel corpus for English-Japanese, for which the amount of publicly available parallel corpora is still limited.
6 papers · 0 benchmarks
JerichoWorld is a dataset that enables the creation of learning agents that can build knowledge graph-based world models of interactive narratives.
6 papers · 2 benchmarks
KAMEL (Knowledge Analysis with Multitoken Entities in Language Models)
KAMEL comprises knowledge about 234 relations from Wikidata with a large training, validation, and test dataset.
6 papers · 1 benchmark
LIS (low-light instance segmentation)
To reveal and systematically investigate the effectiveness of the proposed method in the real world, a real low-light image dataset for instance segmentation is necessary and urgently needed.
6 papers · 0 benchmarks
LSSED, a challenging large-scale english dataset for speech emotion recognition.
6 papers · 1 benchmark
LiDAR-MOS (LiDAR-based Moving Object Segmentation)
Tasks.
6 papers · 0 benchmarks
Libri-adhoc40 is a synchronized speech corpus which collects the replayed Librispeech data from loudspeakers by ad-hoc microphone arrays of 40 strongly synchronized distributed nodes in a real office environment.
6 papers · 0 benchmarks
Lyra is a dataset for code generation that consists on Python code with embedded SQL.
6 papers · 0 benchmarks
MARIDA (Marine Debris Archive)
MARIDA (Marine Debris Archive) is the first dataset based on the multispectral Sentinel-2 (S2) satellite data, which distinguishes Marine Debris from various marine features that co-exist, including Sargassum macroalgae, Ships, Natural…
6 papers · 1 benchmark
MCXFACE (Multi-Channel Heterogeneous Face Recognition dataset)
MCXFace is a heterogeneous face recognition dataset consisting of multi-channel image samples for 51 subjects.
6 papers · 0 benchmarks
The MEDIA French corpus is dedicated to semantic extraction from speech in a context of human/machine dialogues.
6 papers · 0 benchmarks
MMPTRACK (Multi-camera Multiple People Tracking Dataset)
Multi-camera Multiple People Tracking (MMPTRACK) dataset has about 9.6 hours of videos, with over half a million frame-wise annotations.
6 papers · 1 benchmark
The Malimg Dataset contains 9,339 malware byteplot images from 25 different families.
6 papers · 1 benchmark
Generate high-quality 3D ground-truth shapes for reconstruction evaluation is extremely challenging because even 3D scanners can only generate pseudo ground-truth shapes with artefacts.
6 papers · 0 benchmarks
Modern Office-31 is a refurbished version of the commonly used Office-31 dataset.
6 papers · 0 benchmarks
MuChoMusic is a benchmark designed to evaluate music understanding in multimodal language models focused on audio.
6 papers · 0 benchmarks
NASA C-MAPSS-2 (Turbofan Engine Degradation Simulation Data Set-2)
The generation of data-driven prognostics models requires the availability of datasets with run-to-failure trajectories.
6 papers · 1 benchmark
NLU++ (NLLU++ : A Multi-Label, Slot-Rich, Generalisable Dataset for Natural Language Understanding in Task-Oriented Dialogue)
nlu++ is a dataset for natural language understanding (NLU) in task-oriented dialogue (ToD) systems, with the aim to provide a much more challenging evaluation environment for dialogue NLU models, up to date with the current application…
6 papers · 0 benchmarks
There are two versions of the NLmaps corpus.
6 papers · 0 benchmarks
Naamapadam is a Named Entity Recognition (NER) dataset for the 11 major Indian languages from two language families.
6 papers · 0 benchmarks
OTTers is a dataset of human one-turn topic transitions.
6 papers · 0 benchmarks
Open Images is a computer vision dataset covering ~9 million images with labels spanning thousands of object categories.
6 papers · 0 benchmarks
PLABA (Plain Language Adaptation of Biomedical Abstracts)
Plain Language Adaptation of Biomedical Abstracts (PLABA) is a dataset designed for automatic adaptation that is both document- and sentence-aligned.
6 papers · 0 benchmarks
PhysioNet Challenge 2021 (The PhysioNet/Computing in Cardiology Challenge 2021)
Data Description The training data contains twelve-lead ECGs.
6 papers · 2 benchmarks
Prophesee GEN4 Dataset (Prophesee 1 Megapixel Automotive Detection Dataset)
The dataset is split between train, test and val folders.
6 papers · 0 benchmarks
QA-SRL Bank 2.0 is a large-scale corpus of Question-Answer driven Semantic Role Labeling (QA-SRL) annotations.
6 papers · 0 benchmarks
The RAD-ChestCT dataset is a large medical imaging dataset developed by Duke MD/PhD Rachel Draelos during her Computer Science PhD supervised by Lawrence Carin.
6 papers · 0 benchmarks
RELX is a benchmark dataset for cross-lingual relation classification in English, French, German, Spanish and Turkish.
6 papers · 0 benchmarks
RealCQA Scientific Chart Question Answering as a Test-bed for First-Order Logic check on huggingface : https://huggingface.co/datasets/sal4ahm/RealCQA
6 papers · 1 benchmark
SDWPF (A Dataset for Spatial Dynamic Wind Power Forecasting Challenge at KDD Cup 2022)
The unique Spatial Dynamic Wind Power Forecasting dataset: SDWPF, which includes the spatial distribution of wind turbines, as well as the dynamic context factors.
6 papers · 0 benchmarks
SMART-101 (Simple Multimodal Algorithmic Reasoning Task Dataset)
Recent times have witnessed an increasing number of applications of deep neural networks towards solving tasks that require superior cognitive abilities, e.g., playing Go, generating art, ChatGPT, etc.
6 papers · 0 benchmarks
SemEval 2014 is a collection of datasets used for the Semantic Evaluation (SemEval) workshop, an annual event that focuses on the evaluation and comparison of systems that can analyze diverse semantic phenomena in text.
6 papers · 0 benchmarks
A multimodal dataset for sentiment analysis on internet memes.
6 papers · 0 benchmarks
SemOpenAlex is an extensive RDF knowledge graph that contains over 26 billion triples about scientific publications and their associated entities, such as authors, institutions, journals, and concepts.
6 papers · 0 benchmarks
ShapeTalk (The ShapeTalk Dataset)
ShapeTalk contains over half a million discriminative utterances produced by contrasting the shapes of common 3D objects for a variety of object classes and degrees of similarity.
6 papers · 0 benchmarks
ShellcodeIA32 is a dataset containing 20 years of shellcodes from a variety of sources is the largest collection of shellcodes in assembly available to date.
6 papers · 1 benchmark
SherLIiC is a testbed for lexical inference in context (LIiC), consisting of 3985 manually annotated inference rule candidates (InfCands), accompanied by (i) ~960k unlabeled InfCands, and (ii) ~190k typed textual relations between Freebase…
6 papers · 0 benchmarks
ShipSG (Ship Segmentation and Georeferencing Dataset)
The ShipSG dataset is the first public dataset of its kind for ship segmentation and georeferencing.
6 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.