Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 25 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 1153–1200 of 3,998

SimpleQuestionsWikidata maps SimpleQuestions to Wikidata.
6 papers · 1 benchmark
The data and audio included here were collected for the Soundscape Attributes Translation Project (SATP).
6 papers · 0 benchmarks
Swords (Stanford Word Substitution benchmark)
Swords (Standford Word Substitution) is a benchmark for lexical substitution, the task of finding appropriate substitutes for a target word in a context.
6 papers · 0 benchmarks
TREC-10 (TREC-10 Question Classification)
A question type classification dataset with 6 classes for questions about a person, location, numeric information, etc.
6 papers · 1 benchmark
UDD is an underwater open-sea farm object detection dataset.
6 papers · 0 benchmarks
We present a further analysis of visual modality incompleteness, benchmarking latest MMEA models on our proposed dataset MMEA-UMVM.
6 papers · 3 benchmarks
Ultra-high definition benchmark (UHDBench) includes 2293 images at 2k resolution sourced from the ground-truth test sets of HRSOD, LIU4k, UAVid, UHDM, and UHRSD.
6 papers · 1 benchmark
VEDAI (Vehicle Detection in Aerial Imagery)
VEDAI is a dataset for Vehicle Detection in Aerial Imagery, provided as a tool to benchmark automatic target recognition algorithms in unconstrained environments.
6 papers · 1 benchmark
VideoCube is a high-quality and large-scale benchmark to create a challenging real-world experimental environment for Global Instance Tracking (GIT).
6 papers · 1 benchmark
WDC Products is an entity matching benchmark which provides for the systematic evaluation of matching systems along combinations of three dimensions while relying on real-word data.
6 papers · 4 benchmarks
WSJ0-2mix-extr is a speech extraction dataset
6 papers · 1 benchmark
Multi-level Benchmark of Watermarks for Large Language Models
6 papers · 0 benchmarks
WebLINX (Real-World Website Navigation with Multi-Turn)
WebLINX is a large-scale benchmark of 100K interactions across 2300 expert demonstrations of conversational web navigation.
6 papers · 1 benchmark
WikiNEuRal is a high-quality automatically-generated dataset for Multilingual Named Entity Recognition.
6 papers · 0 benchmarks
XImageNet-12 (XIMAGENET-12: An Explainable AI Benchmark Dataset for Model Robustness Evaluation)
Enlarge the dataset to understand how image background effect the Computer Vision ML model.
6 papers · 1 benchmark
XL-BEL is a benchmark for cross-lingual biomedical entity linking (XL-BEL).
6 papers · 0 benchmarks
A new dataset with significant occlusions related to object manipulation.
6 papers · 0 benchmarks
Video object segmentation has been studied extensively in the past decade due to its importance in understanding video spatial-temporal structures as well as its value in industrial applications.
6 papers · 1 benchmark
The ZS-F-VQA dataset is a new split of the F-VQA dataset for zero-shot problem.
6 papers · 1 benchmark
Benchmark dataset for abstracts and titles of 100,000 ArXiv scientific papers.
6 papers · 1 benchmark
exiD Dataset (The Entries and Exits Drone Dataset)
The exiD dataset introduces a groundbreaking collection of naturalistic road user trajectories at highway entries and exits in Germany, meticulously captured with drones to navigate past the limitations of conventional traffic data…
6 papers · 0 benchmarks
tdcommons (Therapeutics Data Commons)
Therapeutics Data Commons is an open-science initiative with AI/ML-ready datasets and AI/ML tasks for therapeutics, spanning the discovery and development of safe and effective medicines.
6 papers · 1 benchmark
1QIsaa data collection (1QIsaa data collection (binarized images, feature files, and plotting scripts) for writer identification test)
This data set is collected for the ERC project: The Hands that Wrote the Bible: Digital Palaeography and Scribal Culture of the Dead Sea Scrolls PI: Mladen Popović Grant agreement ID: 640497 Project website:…
5 papers · 0 benchmarks
25kTrees (Individual Tree Crown Annotations)
Manual crown delineation of individual trees in two countries: Denmark and Finland.
5 papers · 0 benchmarks
2DeteCT (2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning)
Maximilian B.
5 papers · 0 benchmarks
The 2017 PhysioNet/CinC Challenge aims to encourage the development of algorithms to classify, from a single short ECG lead recording (between 30 s and 60 s in length), whether the recording shows normal sinus rhythm, atrial fibrillation…
5 papers · 0 benchmarks
AMALGUM (A Machine Annotated Lookalike of GUM)
AMALGUM is a machine annotated multilayer corpus following the same design and annotation layers as GUM, but substantially larger (around 4M tokens).
5 papers · 0 benchmarks
Acappella comprises around 46 hours of a cappella solo singing videos sourced from YouTbe, sampled across different singers and languages.
5 papers · 0 benchmarks
BB-norm-habitat (Bacteria Biotope - entity normalization - bacterial habitat)
In the BB-norm modality of this task, participant systems had to normalize textual entity mentions according to the OntoBiotope ontology for habitats.
5 papers · 0 benchmarks
BB-norm-phenotype (Bacteria Biotope - entity normalization - phenotype)
In the BB-norm modality of this task, participant systems had to normalize textual entity mentions according to the OntoBiotope ontology for phenotypes.
5 papers · 0 benchmarks
BRACE (The Breakdancing Competition Dataset for Dance Motion Synthesis)
BRACE is a dataset for audio-conditioned dance motion synthesis challenging common assumptions for this task: - strong music-dance correlation - controlled motion data - simple poses and movements To address these issues: - We focus on…
5 papers · 2 benchmarks
BTS3.1 (Expanding Accurate Person Recognition to New Altitudes and Ranges: The BRIAR Dataset)
Large, multimodal biometric dataset: It contains still images and videos of over 1,000 people captured at various ranges (up to 1,000 meters) and elevations (up to 400 meters) using a diverse set of cameras (commercial, military-grade,…
5 papers · 2 benchmarks
BioNLI (Biomedical Natural Language Inference)
BioNLI is a dataset in biomedical natural language inference.
5 papers · 1 benchmark
Blizzard Challenge 2013 (Blizzard Challenge 2013 - English language tasks)
The English data for voice building was obtained, prepared and provided the the challenge by Lessac Technologies Inc., having originally came from the publishers Voice Factory International Inc.
5 papers · 1 benchmark
CBC (Complete Blood Count)
The complete blood count (CBC) dataset contains 360 blood smear images along with their annotation files splitting into Training, Testing, and Validation sets.
5 papers · 0 benchmarks
The dataset offers tag and mask annotations for image-text pairs from the CC3M validation set.
5 papers · 2 benchmarks
CLIP (CLIP: A Dataset for Extracting Action Items for Physicians from Hospital Discharge Notes)
We created a dataset of clinical action items annotated over MIMIC-III.
5 papers · 0 benchmarks
CSFCube is an expert annotated test collection to evaluate models trained to perform faceted Query by Example.
5 papers · 0 benchmarks
ChronoMagic with 2265 metamorphic time-lapse videos, each accompanied by a detailed caption.
5 papers · 0 benchmarks
Coached Conversational Preference Elicitation is a dataset consisting of 502 English dialogs with 12,000 annotated utterances between a user and an assistant discussing movie preferences in natural language.
5 papers · 0 benchmarks
ComFact is a benchmark for commonsense fact linking, where models are given contexts and trained to identify situationally-relevant commonsense knowledge from KGs.
5 papers · 0 benchmarks
CovidET (Emotions and their Triggers during Covid-19)
Crises such as the COVID-19 pandemic continuously threaten our world and emotionally affect billions of people worldwide in distinct ways.
5 papers · 0 benchmarks
DREAM-dataset (Deep Robot-to-camera Extrinsics for Articulated Manipulators)
The DREAM dataset is introduce by the paper "Camera-to-Robot Pose Estimation from a Single Image" (ICRA 2020).
5 papers · 1 benchmark
DeToxy (DeToxy: A Large-Scale Multimodal Dataset for Toxicity Classification in Spoken Utterances)
DeToxy is a publicly available toxicity annotated dataset for the English language.
5 papers · 0 benchmarks
Diabetes (Diabetes 130-US Hospitals for Years 1999-2008)
What do the instances in this dataset represent?
5 papers · 3 benchmarks
The Distress Analysis Interview Corpus/Wizard-of-Oz set (DAIC-WOZ) dataset [24, 25] comprises voice and text samples from 189 interviewed healthy and control persons and their PHQ-8 depression detection questionnaire.
5 papers · 0 benchmarks
DurLAR (A High-Fidelity 128-Channel LiDAR Dataset with Panoramic Ambient and Reflectivity Imagery)
DurLAR is a high-fidelity 128-channel 3D LiDAR dataset with panoramic ambient (near infrared) and reflectivity imagery for multi-modal autonomous driving applications.
5 papers · 0 benchmarks
The EDT dataset is designed for corporate event detection and text-based stock prediction (trading strategy) benchmark.
5 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.