Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 37 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 1729–1776 of 3,998
3DYoga90 (3DYoga90: A Hierarchical Video Dataset for Yoga Pose Understanding)
3DYoga90 is organized within a three-level label hierarchy.
2 papers · 0 benchmarks
We crawled 5000 paper, slide pairs from conference proceeding websites.
2 papers · 0 benchmarks
ADORE (A benchmark dataset for machine learning in ecotoxicology)
ADORE is a benchmark dataset for machine learning for ecotixicology, covering acute aquatic toxicity in three relevant taxonomic groups (fish, crustaceans, and algae).
2 papers · 1 benchmark
Antonio Gulli’s corpus of news articles is a collection of more than 1 million news articles.
2 papers · 0 benchmarks
AIDA/testc is a new challenging test set for entity linking systems containing 131 Reuters news articles published between December 5th and 7th, 2020.
2 papers · 1 benchmark
AIME (AI Music Evaluation Dataset)
The AIME dataset contains 6,000 audio tracks generated by 12 music generation models in addition to 500 tracks from MTG-Jamendo.
2 papers · 0 benchmarks
AM2iCo (Adversarial and Multilingual Meaning in Context)
AM2iCo is a wide-coverage and carefully designed cross-lingual and multilingual evaluation set.
2 papers · 0 benchmarks
ANETAC (Arabic Named Entity Transliteration and Classification)
An English-Arabic named entity transliteration and classification dataset built from freely available parallel translation corpora.
2 papers · 0 benchmarks
APE (Automatic Post-Editing)
APE is useful to evaluate Machine Translation automatic post-editing (APE), which is the task of improving the output of a blackbox MT system by automatically fixing its mistakes.
2 papers · 0 benchmarks
Assessing the value of energy efficiency improvements can be challenging as there's no way to truly know how much energy a building would have used without the improvements.
2 papers · 0 benchmarks
AV Digits Database is an audiovisual database which contains normal, whispered and silent speech.
2 papers · 0 benchmarks
AVCAffe (A Large Scale Audio-Visual Dataset of Cognitive Load and Affect for Remote Work)
We introduce AVCAffe, the first Audio-Visual dataset consisting of Cognitive load and Affect attributes.
2 papers · 0 benchmarks
AVMIT (Audiovisual Moments in Time)
Audiovisual Moments in Time (AVMIT) is a large-scale dataset of audiovisual action events.
2 papers · 0 benchmarks
AVSync15 is a high-quality synchronized audio-video dataset curated from VGGSound.
2 papers · 0 benchmarks
AWARE (AWARE: Aspect-Based Sentiment Analysis Dataset of Apps Reviews for Requirements Elicitation)
The peer-reviewed paper of AWARE dataset is published in ASEW 2021, and can be accessed through: http://doi.org/10.1109/ASEW52652.2021.00049.
2 papers · 3 benchmarks
We present the AWS documentation corpus, an open-book QA dataset, which contains 25,175 documents along with 100 matched questions and answers.
2 papers · 0 benchmarks
ActionBench contains two carefully designed probing tasks: Action Antonym and Video Reversal, which targets multimodal alignment capabilities and temporal understanding skills of the model, respectively.
2 papers · 0 benchmarks
Measurement data related to the publication „Active TLS Stack Fingerprinting: Characterizing TLS Server Deployments at Scale“.
2 papers · 0 benchmarks
Amazon-PQA is a product question-answer dataset.
2 papers · 0 benchmarks
Ambiguous-HOI is a challenging dataset containing ambiguous human-object interaction images for HOI detection based on HICO-DET.
2 papers · 0 benchmarks
About Dataset Step right up to our AI data collection company, where we’ve got something special just for you: a unique set of American Sign Language datasets!
2 papers · 0 benchmarks
Amharic - English Parallel Corpus for Machine Translation contains 33,955 sentence pairs extracted text from such news platforms as Ethiopian Press Agency1, Fana Broadcasting Corporate2, and Walta Information Center3.
2 papers · 0 benchmarks
Our trajectory dataset consists of camera-based images, LiDAR scanned point clouds, and manually annotated trajectories.
2 papers · 1 benchmark
A real-world dataset, with hyper-accurate digital counterpart & comprehensive ground-truth annotation.
2 papers · 1 benchmark
The ArxivPapers dataset is an unlabelled collection of over 104K papers related to machine learning and published on arXiv.org between 2007–2020.
2 papers · 0 benchmarks
The availability of high-quality datasets play a crucial role in advancing research and development especially, for safety critical and autonomous systems.
2 papers · 0 benchmarks
Audio-alpaca: A preference dataset for aligning text-to-audio models Audio-alpaca is a pairwise preference dataset containing about 15k (prompt,chosen, rejected) triplets where given a textual prompt, chosen is the preferred generated…
2 papers · 0 benchmarks
BASEPROD (The Bardenas Semi-Desert Planetary Rover Dataset)
BASEPROD provides comprehensive rover sensor data collected over a 1.7 km traverse, accompanied by high-resolution 2D and 3D drone maps of the terrain.
2 papers · 0 benchmarks
Full-text chemical identification and indexing in PubMed articles.
2 papers · 3 benchmarks
BEE23 (Multi-bee Tracking Benchmark)
We collected 32 videos that record bee colony activity from different periods on several sunny days.
2 papers · 0 benchmarks
BabySLM is a language-acquisition-friendly benchmark to probe speech-based LMs at the lexical and syntactic levels, both of which are compatible with the vocabulary typical of children's language experiences.
2 papers · 0 benchmarks
Dataset of the Beacon3D benchmark: Unveiling the Mist over 3D Vision-Language Understanding: Object-centric Evaluation with Chain-of-Analysis.
2 papers · 0 benchmarks
BiGe (Bielefeld Gesture Corpus)
The BiGe corpus is comprised of 54.360 shots of interest extracted from TED and TEDx talks.
2 papers · 0 benchmarks
BiasCorp is a dataset for racism detection containing 139,090 comments and news segment from three specific sources - Fox News, BreitbartNews and YouTube.
2 papers · 0 benchmarks
BioVid (BioVid Heat Pain Database)
To advance methods for pain assessment, in particular automatic assessment methods, the BioVid Heat Pain Database was collected in a collaboration of the Neuro-Information Technology group of the University of Magdeburg and the Medical…
2 papers · 0 benchmarks
Biographical (Biographical: A Semi-Supervised Relation Extraction Dataset)
Biographical is a semi-supervised dataset for RE.
2 papers · 0 benchmarks
Human keypoint dataset of anime/manga-style character illustrations.
2 papers · 0 benchmarks
BlendMimic3D (A Synthetic Dataset for Human Pose Estimation)
BlendMimic3D is a pioneering synthetic dataset developed using Blender, designed to enhance Human Pose Estimation (HPE) research.
2 papers · 0 benchmarks
Dataset created in the paper "Learning to Count Objects in Images" by Victor Lempitsky and Andrew Zisserman exists as a benchmark to have a dataset useful for cell enumeration.
2 papers · 0 benchmarks
BoostCLIR is a bilingual (Japanese-English) corpus of patent abstracts, extracted from the MAREC patent data, and the data from the NTCIR PatentMT workshop collections, accompanied with relevance judgements for the task of patent prior-art…
2 papers · 0 benchmarks
Several datasets are fostering innovation in higher-level functions for everyone, everywhere.
2 papers · 0 benchmarks
BugRepo maintains a collection of bug reports that are publicly available for research purposes.
2 papers · 0 benchmarks
C2A: Combination to Application Dataset Overview This repository contains the code and information for the paper "UAV-Enhanced Combination to Application: Comprehensive Analysis and Benchmarking of a Human Detection Dataset for Disaster…
2 papers · 1 benchmark
CAP (Consented Activities of People)
The Consented Activities of People (CAP) dataset is a fine grained activity dataset for visual AI research curated using the Visym Collector platform.
2 papers · 0 benchmarks
CAR (Cityscapes Attributes Recognition)
CAR contains visual attributes for objects in the Cityscapes dataset.
2 papers · 0 benchmarks
CARBEN (Composite Adversarial Robustness Benchmark)
Prior literature on adversarial attack methods has mainly focused on attacking with and defending against a single threat model, e.g., perturbations bounded in Lp ball.
2 papers · 0 benchmarks
CAVES (A Dataset to facilitate Explainable Classification and Summarization of Concerns towards COVID Vaccines)
CAVES is the first large-scale dataset containing about 10k COVID-19 anti-vaccine tweets labelled into various specific anti-vaccine concerns in a multi-label setting.
2 papers · 0 benchmarks
CommonCrawl News is a dataset containing news articles from news sites all over the world.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.