Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 25 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 1153–1200 of 3,998
SimpleQuestionsWikidata maps SimpleQuestions to Wikidata.
6 papers · 1 benchmark
The data and audio included here were collected for the Soundscape Attributes Translation Project (SATP).
6 papers · 0 benchmarks
Swords (Stanford Word Substitution benchmark)
Swords (Standford Word Substitution) is a benchmark for lexical substitution, the task of finding appropriate substitutes for a target word in a context.
6 papers · 0 benchmarks
TREC-10 (TREC-10 Question Classification)
A question type classification dataset with 6 classes for questions about a person, location, numeric information, etc.
6 papers · 1 benchmark
UDD is an underwater open-sea farm object detection dataset.
6 papers · 0 benchmarks
We present a further analysis of visual modality incompleteness, benchmarking latest MMEA models on our proposed dataset MMEA-UMVM.
6 papers · 3 benchmarks
Ultra-high definition benchmark (UHDBench) includes 2293 images at 2k resolution sourced from the ground-truth test sets of HRSOD, LIU4k, UAVid, UHDM, and UHRSD.
6 papers · 1 benchmark
VEDAI (Vehicle Detection in Aerial Imagery)
VEDAI is a dataset for Vehicle Detection in Aerial Imagery, provided as a tool to benchmark automatic target recognition algorithms in unconstrained environments.
6 papers · 1 benchmark
VideoCube is a high-quality and large-scale benchmark to create a challenging real-world experimental environment for Global Instance Tracking (GIT).
6 papers · 1 benchmark
WDC Products is an entity matching benchmark which provides for the systematic evaluation of matching systems along combinations of three dimensions while relying on real-word data.
6 papers · 4 benchmarks
WSJ0-2mix-extr is a speech extraction dataset
6 papers · 1 benchmark
Multi-level Benchmark of Watermarks for Large Language Models
6 papers · 0 benchmarks
WebLINX (Real-World Website Navigation with Multi-Turn)
WebLINX is a large-scale benchmark of 100K interactions across 2300 expert demonstrations of conversational web navigation.
6 papers · 1 benchmark
WikiNEuRal is a high-quality automatically-generated dataset for Multilingual Named Entity Recognition.
6 papers · 0 benchmarks
XImageNet-12 (XIMAGENET-12: An Explainable AI Benchmark Dataset for Model Robustness Evaluation)
Enlarge the dataset to understand how image background effect the Computer Vision ML model.
6 papers · 1 benchmark
XL-BEL is a benchmark for cross-lingual biomedical entity linking (XL-BEL).
6 papers · 0 benchmarks
A new dataset with significant occlusions related to object manipulation.
6 papers · 0 benchmarks
Video object segmentation has been studied extensively in the past decade due to its importance in understanding video spatial-temporal structures as well as its value in industrial applications.
6 papers · 1 benchmark
The ZS-F-VQA dataset is a new split of the F-VQA dataset for zero-shot problem.
6 papers · 1 benchmark
Benchmark dataset for abstracts and titles of 100,000 ArXiv scientific papers.
6 papers · 1 benchmark
The exiD dataset introduces a groundbreaking collection of naturalistic road user trajectories at highway entries and exits in Germany, meticulously captured with drones to navigate past the limitations of conventional traffic data…
6 papers · 0 benchmarks
Therapeutics Data Commons is an open-science initiative with AI/ML-ready datasets and AI/ML tasks for therapeutics, spanning the discovery and development of safe and effective medicines.
6 papers · 1 benchmark
1QIsaa data collection (1QIsaa data collection (binarized images, feature files, and plotting scripts) for writer identification test)
This data set is collected for the ERC project: The Hands that Wrote the Bible: Digital Palaeography and Scribal Culture of the Dead Sea Scrolls PI: Mladen Popović Grant agreement ID: 640497 Project website:…
5 papers · 0 benchmarks
25kTrees (Individual Tree Crown Annotations)
Manual crown delineation of individual trees in two countries: Denmark and Finland.
5 papers · 0 benchmarks
2DeteCT (2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning)
Maximilian B.
5 papers · 0 benchmarks
The 2017 PhysioNet/CinC Challenge aims to encourage the development of algorithms to classify, from a single short ECG lead recording (between 30 s and 60 s in length), whether the recording shows normal sinus rhythm, atrial fibrillation…
5 papers · 0 benchmarks
AMALGUM (A Machine Annotated Lookalike of GUM)
AMALGUM is a machine annotated multilayer corpus following the same design and annotation layers as GUM, but substantially larger (around 4M tokens).
5 papers · 0 benchmarks
Acappella comprises around 46 hours of a cappella solo singing videos sourced from YouTbe, sampled across different singers and languages.
5 papers · 0 benchmarks
BB-norm-habitat (Bacteria Biotope - entity normalization - bacterial habitat)
In the BB-norm modality of this task, participant systems had to normalize textual entity mentions according to the OntoBiotope ontology for habitats.
5 papers · 0 benchmarks
In the BB-norm modality of this task, participant systems had to normalize textual entity mentions according to the OntoBiotope ontology for phenotypes.
5 papers · 0 benchmarks
BRACE (The Breakdancing Competition Dataset for Dance Motion Synthesis)
BRACE is a dataset for audio-conditioned dance motion synthesis challenging common assumptions for this task: - strong music-dance correlation - controlled motion data - simple poses and movements To address these issues: - We focus on…
5 papers · 2 benchmarks
BTS3.1 (Expanding Accurate Person Recognition to New Altitudes and Ranges: The BRIAR Dataset)
Large, multimodal biometric dataset: It contains still images and videos of over 1,000 people captured at various ranges (up to 1,000 meters) and elevations (up to 400 meters) using a diverse set of cameras (commercial, military-grade,…
5 papers · 2 benchmarks
BioNLI (Biomedical Natural Language Inference)
BioNLI is a dataset in biomedical natural language inference.
5 papers · 1 benchmark
The English data for voice building was obtained, prepared and provided the the challenge by Lessac Technologies Inc., having originally came from the publishers Voice Factory International Inc.
5 papers · 1 benchmark
CBC (Complete Blood Count)
The complete blood count (CBC) dataset contains 360 blood smear images along with their annotation files splitting into Training, Testing, and Validation sets.
5 papers · 0 benchmarks
The dataset offers tag and mask annotations for image-text pairs from the CC3M validation set.
5 papers · 2 benchmarks
CLIP (CLIP: A Dataset for Extracting Action Items for Physicians from Hospital Discharge Notes)
We created a dataset of clinical action items annotated over MIMIC-III.
5 papers · 0 benchmarks
CSFCube is an expert annotated test collection to evaluate models trained to perform faceted Query by Example.
5 papers · 0 benchmarks
ChronoMagic with 2265 metamorphic time-lapse videos, each accompanied by a detailed caption.
5 papers · 0 benchmarks
Coached Conversational Preference Elicitation is a dataset consisting of 502 English dialogs with 12,000 annotated utterances between a user and an assistant discussing movie preferences in natural language.
5 papers · 0 benchmarks
ComFact is a benchmark for commonsense fact linking, where models are given contexts and trained to identify situationally-relevant commonsense knowledge from KGs.
5 papers · 0 benchmarks
CovidET (Emotions and their Triggers during Covid-19)
Crises such as the COVID-19 pandemic continuously threaten our world and emotionally affect billions of people worldwide in distinct ways.
5 papers · 0 benchmarks
DREAM-dataset (Deep Robot-to-camera Extrinsics for Articulated Manipulators)
The DREAM dataset is introduce by the paper "Camera-to-Robot Pose Estimation from a Single Image" (ICRA 2020).
5 papers · 1 benchmark
DeToxy (DeToxy: A Large-Scale Multimodal Dataset for Toxicity Classification in Spoken Utterances)
DeToxy is a publicly available toxicity annotated dataset for the English language.
5 papers · 0 benchmarks
Diabetes (Diabetes 130-US Hospitals for Years 1999-2008)
What do the instances in this dataset represent?
5 papers · 3 benchmarks
The Distress Analysis Interview Corpus/Wizard-of-Oz set (DAIC-WOZ) dataset [24, 25] comprises voice and text samples from 189 interviewed healthy and control persons and their PHQ-8 depression detection questionnaire.
5 papers · 0 benchmarks
DurLAR (A High-Fidelity 128-Channel LiDAR Dataset with Panoramic Ambient and Reflectivity Imagery)
DurLAR is a high-fidelity 128-channel 3D LiDAR dataset with panoramic ambient (near infrared) and reflectivity imagery for multi-modal autonomous driving applications.
5 papers · 0 benchmarks
The EDT dataset is designed for corporate event detection and text-based stock prediction (trading strategy) benchmark.
5 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.