Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 22 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 1009–1056 of 3,998
WorldStrat (The WorldStrat Dataset: Open High-Resolution Satellite Imagery With Paired Multi-Temporal Low-Resolution)
Nearly 10,000 km² of free high-resolution and paired multi-temporal low-resolution satellite imagery of unique locations which ensure stratified representation of all types of land-use across the world: from agriculture to ice caps, from…
8 papers · 0 benchmarks
X3D is a dataset containing 15 scenes and covering 4 applications for X-ray 3D reconstruction.
8 papers · 2 benchmarks
YouTube-ASL is a large-scale, open-domain corpus of American Sign Language (ASL) videos and accompanying English captions drawn from YouTube.
8 papers · 0 benchmarks
We address the problem of automatically learning the main steps to complete a certain task, such as changing a car tire, from a set of narrated instruction videos.
8 papers · 2 benchmarks
ZEB (Zero-shot Evaluation Benchmark)
A evaluation benchmark ZEB for image matching by merging 8 real-world datasets and 4 simulated datasets with diverse image resolutions, scene conditions and view points.
8 papers · 1 benchmark
3DOH50K is the first real 3D human dataset for the problem of human reconstruction and pose estimation in occlusion scenarios.
7 papers · 1 benchmark
Aachen Day-Night v1.1 dataset is an extended version of the original Aachen Day-Night dataset.
7 papers · 1 benchmark
We introduce ArtBench-10, the first class-balanced, high-quality, cleanly annotated, and standardized dataset for benchmarking artwork generation.
7 papers · 1 benchmark
ArtiFact (Artificial and Factual Image Dataset for Synthetic Image Detection)
The ArtiFact dataset is a large-scale image dataset that aims to include a diverse collection of real and synthetic images from multiple categories, including Human/Human Faces, Animal/Animal Faces, Places, Vehicles, Art, and many other…
7 papers · 0 benchmarks
Prediction of Finger Flexion IV Brain-Computer Interface Data Competition The goal of this dataset is to predict the flexion of individual fingers from signals recorded from the surface of the brain (electrocorticography (ECoG)).
7 papers · 1 benchmark
BiSECT is a dataset for sentence simplification, which is the ability to take a long, complex sentence and split it into shorter sentences, rephrasing as necessary.
7 papers · 0 benchmarks
CEFR-SP contains 17k English sentences annotated with the levels based on the Common European Framework of Reference for Languages assigned by English-education professionals.
7 papers · 0 benchmarks
ChangeSim is a dataset aimed at online scene change detection (SCD) and more.
7 papers · 2 benchmarks
The CheXmask Database presents a comprehensive, uniformly annotated collection of chest radiographs, constructed from five public databases: ChestX-ray8, Chexpert, MIMIC-CXR-JPG, Padchest and VinDr-CXR.
7 papers · 0 benchmarks
The dataset published here is the largest, most diverse and consistent crack segmentation dataset constructed so far.
7 papers · 0 benchmarks
In this work we create a question answering dataset over the DBLP scholarly knowledge graph (KG).
7 papers · 0 benchmarks
DOLPHINS (Dataset for Collaborative Perception enabled Harmonious and Interconnected Self-driving)
Vehicle-to-Everything (V2X) network has enabled collaborative perception in autonomous driving, which is a promising solution to the fundamental defect of stand-alone intelligence including blind zones and long-range perception.
7 papers · 0 benchmarks
Evaluate a natural language code generation model on real data science pedagogical notebooks!
7 papers · 0 benchmarks
Demetr is a diagnostic dataset with 31K English examples (translated from 10 source languages) for evaluating the sensitivity of MT evaluation metrics to 35 different linguistic perturbations spanning semantic, syntactic, and morphological…
7 papers · 0 benchmarks
FOR-instance (FOR-instance: a UAV laser scanning benchmark dataset for semantic and instance segmentation of individual trees)
The challenge of accurately segmenting individual trees from laser scanning data hinders the assessment of crucial tree parameters necessary for effective forest management, impacting many downstream applications.
7 papers · 0 benchmarks
GeoDE is a geographically diverse dataset with 61,940 images from 40 classes and 6 world regions, and no personally identifiable information, collected through crowd-sourcing.
7 papers · 0 benchmarks
HiREST (HIerarchical REtrieval and STep-captioning)
HiREST (HIerarchical REtrieval and STep-captioning) dataset is a benchmark that covers hierarchical information retrieval and visual/textual stepwise summarization from an instructional video corpus.
7 papers · 0 benchmarks
HyperRED (Hyper-Relational Extraction Dataset)
HyperRED is a dataset for the new task of hyper-relational extraction, which extracts relation triplets together with qualifier information such as time, quantity or location.
7 papers · 1 benchmark
IQUAD (Interactive Question Answering Dataset)
IQUAD is a dataset for Visual Question Answering in interactive environments.
7 papers · 0 benchmarks
KiTS19 (The 2019 Kidney and Kidney Tumor Segmentation Challenge)
The 2021 Kidney and Kidney Tumor Segmentation challenge (abbreviated KiTS21) is a competition in which teams compete to develop the best system for automatic semantic segmentation of renal tumors and surrounding anatomy.
7 papers · 1 benchmark
Language-molecule models have emerged as an exciting direction for molecular discovery and understanding.
7 papers · 1 benchmark
LaFAN1 (Ubisoft La Forge Animation Dataset)
Ubisoft La Forge Animation Dataset ("LAFAN1") Ubisoft La Forge Animation dataset and accompanying code for the SIGGRAPH 2020 paper Robust Motion In-betweening.
7 papers · 1 benchmark
Large Scale Composed Image Retrieval (LaSCo) is a new dataset for Composed Image Retrieval (CoIR), x10 times larger than current ones.
7 papers · 1 benchmark
This work proposes Long-RVOS, a large-scale benchmark for long-term video object segmentation.
7 papers · 1 benchmark
MIPE (Improving Paratope and Epitope Prediction by Multi-Modal Contrastive Learning and Interaction Informativeness Estimation)
Datasets.
7 papers · 1 benchmark
The MMBody dataset provides human body data with motion capture, GT mesh, Kinect RGBD, and millimeter wave sensor data.
7 papers · 0 benchmarks
MSU BASED (MSU BASED Video Deblurring Dataset and Benchmark)
Qualitative dataset with real blurred videos, created by using beam-splitter setup in lab environment
7 papers · 1 benchmark
The Medical Dataset for Abbreviation Disambiguation for Natural Language Understanding (MeDAL) is a large medical text dataset curated for abbreviation disambiguation, designed for natural language understanding pre-training in the medical…
7 papers · 0 benchmarks
In MutualFriends, two agents, A and B, each have a private knowledge base, which contains a list of friends with multiple attributes (e.g., name, school, major, etc.).
7 papers · 0 benchmarks
NASA C-MAPSS (Turbofan Engine Degradation Simulation Data Set)
Engine degradation simulation was carried out using C-MAPSS.
7 papers · 2 benchmarks
NExT-QA is a VideoQA benchmark targeting the explanation of video contents.
7 papers · 1 benchmark
New3, a set of 527 instances from AMR 3.0, whose original source was the LORELEI DARPA project – not included in the AMR 2.0 training set – consisting of excerpts from newswires and online forum.
7 papers · 1 benchmark
OPOSUM is a dataset for the training and evaluation of Opinion Summarization models which contains Amazon reviews from six product domains: Laptop Bags, Bluetooth Headsets, Boots, Keyboards, Televisions, and Vacuums.
7 papers · 0 benchmarks
PDNC (Project Dialogism Novel Corpus)
A annotated dataset of quotations and within-quotation-mentions in 22 full-length English novels.
7 papers · 0 benchmarks
PET (PET: A new Dataset for Process Extraction from Natural Language Text)
The dataset contains 45 documents containing narrative description of business process and their annotations.
7 papers · 0 benchmarks
PHM2017 is a new dataset consisting of 7,192 English tweets across six diseases and conditions: Alzheimer’s Disease, heart attack (any severity), Parkinson’s disease, cancer (any type), Depression (any severity), and Stroke.
7 papers · 0 benchmarks
The PMData dataset aims to combine the traditional lifelogging with sports activity logging.
7 papers · 0 benchmarks
POT-210 (Planar Object Tracking in the Wild: A Benchmark)
Planar object tracking is an actively studied problem in vision-based robotic applications.
7 papers · 0 benchmarks
Electrocardiography (ECG) is a key diagnostic tool to assess the cardiac condition of a patient.
7 papers · 2 benchmarks
ataset format Each row in the dataset splits represents one instance and contains the following tab-separated columns: articleid - article id corresponding to the id of the claim in the LIAR dataset statement - the text of the claim author…
7 papers · 0 benchmarks
These are the protein-ligand complexes of the PoseBusters Benchmark set as described in the paper "PoseBusters: AI-based docking methods fail to generate physically valid poses or generalise to novel sequences" [1] with associated code at…
7 papers · 0 benchmarks
RCooper (Roadside Cooperative Perception Dataset)
The first real-world, large-scale Roadside Cooperative Perception Dataset, RCooper, is released to bloom research on roadside cooperative perception for practical applications.
7 papers · 0 benchmarks
Reddit Corpus is part of a repository of conversational datasets consisting of hundreds of millions of examples, and a standardised evaluation procedure for conversational response selection models using '1-of-100 accuracy'.
7 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.