Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 37 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 1729–1776 of 3,998

3DYoga90 (3DYoga90: A Hierarchical Video Dataset for Yoga Pose Understanding)
3DYoga90 is organized within a three-level label hierarchy.
2 papers · 0 benchmarks
5k_presetation_slides (5000 presentation slide pairs)
We crawled 5000 paper, slide pairs from conference proceeding websites.
2 papers · 0 benchmarks
ADORE (A benchmark dataset for machine learning in ecotoxicology)
ADORE is a benchmark dataset for machine learning for ecotixicology, covering acute aquatic toxicity in three relevant taxonomic groups (fish, crustaceans, and algae).
2 papers · 1 benchmark
AG’s Corpus (AG's corpus of news articlesNews)
Antonio Gulli’s corpus of news articles is a collection of more than 1 million news articles.
2 papers · 0 benchmarks
AIDA/testc is a new challenging test set for entity linking systems containing 131 Reuters news articles published between December 5th and 7th, 2020.
2 papers · 1 benchmark
AIME (AI Music Evaluation Dataset)
The AIME dataset contains 6,000 audio tracks generated by 12 music generation models in addition to 500 tracks from MTG-Jamendo.
2 papers · 0 benchmarks
AM2iCo (Adversarial and Multilingual Meaning in Context)
AM2iCo is a wide-coverage and carefully designed cross-lingual and multilingual evaluation set.
2 papers · 0 benchmarks
ANETAC (Arabic Named Entity Transliteration and Classification)
An English-Arabic named entity transliteration and classification dataset built from freely available parallel translation corpora.
2 papers · 0 benchmarks
APE (Automatic Post-Editing)
APE is useful to evaluate Machine Translation automatic post-editing (APE), which is the task of improving the output of a blackbox MT system by automatically fixing its mistakes.
2 papers · 0 benchmarks
Assessing the value of energy efficiency improvements can be challenging as there's no way to truly know how much energy a building would have used without the improvements.
2 papers · 0 benchmarks
AV Digits Database is an audiovisual database which contains normal, whispered and silent speech.
2 papers · 0 benchmarks
AVCAffe (A Large Scale Audio-Visual Dataset of Cognitive Load and Affect for Remote Work)
We introduce AVCAffe, the first Audio-Visual dataset consisting of Cognitive load and Affect attributes.
2 papers · 0 benchmarks
AVMIT (Audiovisual Moments in Time)
Audiovisual Moments in Time (AVMIT) is a large-scale dataset of audiovisual action events.
2 papers · 0 benchmarks
AVSync15 is a high-quality synchronized audio-video dataset curated from VGGSound.
2 papers · 0 benchmarks
AWARE (AWARE: Aspect-Based Sentiment Analysis Dataset of Apps Reviews for Requirements Elicitation)
The peer-reviewed paper of AWARE dataset is published in ASEW 2021, and can be accessed through: http://doi.org/10.1109/ASEW52652.2021.00049.
2 papers · 3 benchmarks
We present the AWS documentation corpus, an open-book QA dataset, which contains 25,175 documents along with 100 matched questions and answers.
2 papers · 0 benchmarks
ActionBench contains two carefully designed probing tasks: Action Antonym and Video Reversal, which targets multimodal alignment capabilities and temporal understanding skills of the model, respectively.
2 papers · 0 benchmarks
Measurement data related to the publication „Active TLS Stack Fingerprinting: Characterizing TLS Server Deployments at Scale“.
2 papers · 0 benchmarks
Amazon-PQA is a product question-answer dataset.
2 papers · 0 benchmarks
Ambiguous-HOI is a challenging dataset containing ambiguous human-object interaction images for HOI detection based on HICO-DET.
2 papers · 0 benchmarks
About Dataset Step right up to our AI data collection company, where we’ve got something special just for you: a unique set of American Sign Language datasets!
2 papers · 0 benchmarks
Amharic - English Parallel Corpus for Machine Translation contains 33,955 sentence pairs extracted text from such news platforms as Ethiopian Press Agency1, Fana Broadcasting Corporate2, and Walta Information Center3.
2 papers · 0 benchmarks
Our trajectory dataset consists of camera-based images, LiDAR scanned point clouds, and manually annotated trajectories.
2 papers · 1 benchmark
A real-world dataset, with hyper-accurate digital counterpart & comprehensive ground-truth annotation.
2 papers · 1 benchmark
The ArxivPapers dataset is an unlabelled collection of over 104K papers related to machine learning and published on arXiv.org between 2007–2020.
2 papers · 0 benchmarks
The availability of high-quality datasets play a crucial role in advancing research and development especially, for safety critical and autonomous systems.
2 papers · 0 benchmarks
Audio-alpaca: A preference dataset for aligning text-to-audio models Audio-alpaca is a pairwise preference dataset containing about 15k (prompt,chosen, rejected) triplets where given a textual prompt, chosen is the preferred generated…
2 papers · 0 benchmarks
BASEPROD (The Bardenas Semi-Desert Planetary Rover Dataset)
BASEPROD provides comprehensive rover sensor data collected over a 1.7 km traverse, accompanied by high-resolution 2D and 3D drone maps of the terrain.
2 papers · 0 benchmarks
BC7 NLM-Chem (BioCreative VII NLM-Chem)
Full-text chemical identification and indexing in PubMed articles.
2 papers · 3 benchmarks
BEE23 (Multi-bee Tracking Benchmark)
We collected 32 videos that record bee colony activity from different periods on several sunny days.
2 papers · 0 benchmarks
BabySLM is a language-acquisition-friendly benchmark to probe speech-based LMs at the lexical and syntactic levels, both of which are compatible with the vocabulary typical of children's language experiences.
2 papers · 0 benchmarks
Dataset of the Beacon3D benchmark: Unveiling the Mist over 3D Vision-Language Understanding: Object-centric Evaluation with Chain-of-Analysis.
2 papers · 0 benchmarks
BiGe (Bielefeld Gesture Corpus)
The BiGe corpus is comprised of 54.360 shots of interest extracted from TED and TEDx talks.
2 papers · 0 benchmarks
BiasCorp is a dataset for racism detection containing 139,090 comments and news segment from three specific sources - Fox News, BreitbartNews and YouTube.
2 papers · 0 benchmarks
BioVid (BioVid Heat Pain Database)
To advance methods for pain assessment, in particular automatic assessment methods, the BioVid Heat Pain Database was collected in a collaboration of the Neuro-Information Technology group of the University of Magdeburg and the Medical…
2 papers · 0 benchmarks
Biographical (Biographical: A Semi-Supervised Relation Extraction Dataset)
Biographical is a semi-supervised dataset for RE.
2 papers · 0 benchmarks
Bizarre Pose Dataset (Bizarre Pose Dataset of Illustrated Characters)
Human keypoint dataset of anime/manga-style character illustrations.
2 papers · 0 benchmarks
BlendMimic3D (A Synthetic Dataset for Human Pose Estimation)
BlendMimic3D is a pioneering synthetic dataset developed using Blender, designed to enhance Human Pose Estimation (HPE) research.
2 papers · 0 benchmarks
Blue Cells enumeration dataset (Learning to Count Objects in Images blue cells dataset)
Dataset created in the paper "Learning to Count Objects in Images" by Victor Lempitsky and Andrew Zisserman exists as a benchmark to have a dataset useful for cell enumeration.
2 papers · 0 benchmarks
BoostCLIR is a bilingual (Japanese-English) corpus of patent abstracts, extracted from the MAREC patent data, and the data from the NTCIR PatentMT workshop collections, accompanied with relevance judgements for the task of patent prior-art…
2 papers · 0 benchmarks
BreastClassifications4 ([MIMBCD-UI] UTA4: Severity & Pathology Classifications Dataset)
Several datasets are fostering innovation in higher-level functions for everyone, everywhere.
2 papers · 0 benchmarks
BugRepo (Bug Reports)
BugRepo maintains a collection of bug reports that are publicly available for research purposes.
2 papers · 0 benchmarks
C2A: Human Detection in Disaster Scenarios (Combination to Application)
C2A: Combination to Application Dataset Overview This repository contains the code and information for the paper "UAV-Enhanced Combination to Application: Comprehensive Analysis and Benchmarking of a Human Detection Dataset for Disaster…
2 papers · 1 benchmark
CAP (Consented Activities of People)
The Consented Activities of People (CAP) dataset is a fine grained activity dataset for visual AI research curated using the Visym Collector platform.
2 papers · 0 benchmarks
CAR (Cityscapes Attributes Recognition)
CAR contains visual attributes for objects in the Cityscapes dataset.
2 papers · 0 benchmarks
CARBEN (Composite Adversarial Robustness Benchmark)
Prior literature on adversarial attack methods has mainly focused on attacking with and defending against a single threat model, e.g., perturbations bounded in Lp ball.
2 papers · 0 benchmarks
CAVES (A Dataset to facilitate Explainable Classification and Summarization of Concerns towards COVID Vaccines)
CAVES is the first large-scale dataset containing about 10k COVID-19 anti-vaccine tweets labelled into various specific anti-vaccine concerns in a multi-label setting.
2 papers · 0 benchmarks
CC-News (CommonCrawl News dataset)
CommonCrawl News is a dataset containing news articles from news sites all over the world.
2 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.