Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 39 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 1825–1872 of 3,998

This experiment was performed in order to empirically measure the energy use of small, electric Unmanned Aerial Vehicles (UAVs).
2 papers · 1 benchmark
Database of axial impact simulations of the crash box (Database for crashworthiness optimisation)
This repository contains the database of the FEM simulation of axially impacted various configurations of the square crash boxes.
2 papers · 0 benchmarks
This is a part of dataset of the paper published in UAI 2021 (37th Conference on Uncertainty in Artificial Intelligence).
2 papers · 0 benchmarks
DeCOCO is a bilingual (English-German) corpus of image descriptions, where the English part is extracted from the COCO dataset, and the German part are translations by a native German speaker.
2 papers · 0 benchmarks
Deep PCB (Deep Printed Circuit Board)
DeepPCB Dataset Link : A dataset contains 1,500 image pairs, each of which consists of a defect-free template image and an aligned tested image with annotations including positions of 6 most common types of PCB defects: open, short,…
2 papers · 1 benchmark
This collection contains data and code associated with the IPCAI/IJCARS 2020 paper “Automatic Annotation of Hip Anatomy in Fluoroscopy for Robust and Efficient 2D/3D Registration.” The data hosted here consists of annotated datasets of…
2 papers · 0 benchmarks
DenseUAV is a dataset of drone and satellite perspectives collected from 14 universities in low-altitude urban scenes.
2 papers · 0 benchmarks
The Dialogue Fairness dataset is used to evaluate and understand fairness in dialogue models, focusing on gender and racial biases.
2 papers · 0 benchmarks
This is a dataset used to test deep learning-supported deep learning for fault diagnosis: - A digital twin model for a robot.
2 papers · 1 benchmark
DisKnE (Disease Knowledge Evaluation)
DisKnE is a benchmark for Disease Knowledge Evaluation built from MedNLI and MEDIQA-NLI.
2 papers · 0 benchmarks
DropletVideo is a project exploring high-order spatio-temporal consistency in image-to-video generation.
2 papers · 0 benchmarks
Estimating camera motion in deformable scenes poses a complex and open research challenge.
2 papers · 1 benchmark
E-ReDial (Explainable Recommendation Dialogues)
E-ReDial is a conversational recommender system dataset with high-quality explanations.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 1 benchmark
The EARS-Reverb dataset uses real recorded room impulse responses (RIRs) from multiple public datasets (ACE-Challenge, AIR, ARNI, BRUDEX, dEchorate, DetmoldSRIR, and Palimpsest).
2 papers · 1 benchmark
EDGAR-CORPUS is a novel corpus comprising annual reports from all the publicly traded companies in the US spanning a period of more than 25 years.
2 papers · 0 benchmarks
EFO-1-QA is a new dataset to benchmark the combinatorial generalizability of Complex Query Answering (CQA) models by including 301 different queries types, which is 20 times larger than existing datasets.
2 papers · 0 benchmarks
EHE (Elderly Home Exercise)
Human Action Evaluation (HAE) has rarely been applied to real-world disease monitoring, the EHE dataset aims to gather sample data to validate effective HAE methods that could then be expanded on a larger validation scale.
2 papers · 1 benchmark
ENRICH (Multi-purposE dataset for beNchmaRking In Computer vision and pHotogrammetry)
A new synthetic, multi-purpose dataset - called ENRICH - for testing photogrammetric and computer vision algorithms.
2 papers · 0 benchmarks
ENTIGEN (Ethical NaTural Language Interventions in Text-to-Image GENeration)
ENTIGEN is a benchmark dataset to evaluate the change in image generations conditional on ethical interventions across three social axes -- gender, skin color, and culture.
2 papers · 0 benchmarks
ESB (End-to-End Speech Benchmark)
ESB is a benchmark for evaluating the performance of a single automatic speech recognition (ASR) system across a broad set of speech datasets.
2 papers · 0 benchmarks
To automatically generate Python and assembly programs used for security exploits, we curated a large dataset for feeding NMT techniques.
2 papers · 0 benchmarks
The Eedi dataset contains from two school years (September 2018 to May 2020) of students’ answers to mathematics questions from Eedi, a leading educational platform which millions of students interact with daily around the globe.
2 papers · 0 benchmarks
This data was collected by performing a breadth-first search on the user-product-review graph until termination, meaning that it is a fairly comprehensive collection of English-language product data.
2 papers · 1 benchmark
An open corpus of Scientific Research papers which has a representative sample from across scientific disciplines.
2 papers · 0 benchmarks
EmoCause is a dataset of annotated emotion cause words in emotional situations from the EmpatheticDialogues valid and test set.
2 papers · 1 benchmark
The dataset consists of biomedical articles describing randomized control trials (RCTs) that compare multiple treatments.
2 papers · 0 benchmarks
ExHVV is a novel dataset that offers natural language explanations of connotative roles for three types of entities -- heroes, villains, and victims, encompassing 4,680 entities present in 3K memes.
2 papers · 0 benchmarks
ExpMRC is a benchmark for the Explainability evaluation of Machine Reading Comprehension.
2 papers · 0 benchmarks
A new fraud detection dataset FDCompCN for detecting financial statement fraud of companies in China.
2 papers · 1 benchmark
FINDSum (Financial Report Document Summarization)
FINDSum is a large-scale dataset for long text and multi-table summarization.
2 papers · 0 benchmarks
FSVQA (Full-Sentence Visual Question Answering)
Full-Sentence Visual Question Answering (FSVQA) dataset, consisting of nearly 1 million pairs of questions and full-sentence answers for images, built by applying a number of rule-based natural language processing techniques to original…
2 papers · 0 benchmarks
FaceOcc (Face Occlusion Dataset)
FaceOcc is a high-quality face occlusion dataset which contains all mislabeled occlusions in CelebAMask-HQ and complements some occlusions and textures from the internet.
2 papers · 0 benchmarks
FairPrism is a dataset of 5,000 examples of AI-generated English text with detailed human annotations covering a diverse set of harms relating to gender and sexuality.
2 papers · 0 benchmarks
Fashion-MNT is large-scale bilingual product description dataset called Fashion-MMT, which contains over 114k noisy and 40k manually cleaned description translations with multiple product images.
2 papers · 0 benchmarks
We provide multiple human annotations for each test image in Fashion-MNIST.
2 papers · 0 benchmarks
📄 Read 💾 Code 🔗 Webpage 💻 Demo 🤗 Huggingface Dataset 💬 Discussions Overview Users interact with QA systems and leave feedback.
2 papers · 0 benchmarks
Fetoscopic Placental Vessel Segmentation and Registration (FetReg2021) challenge was organized as part of the MICCAI2021 Endoscopic Vision (EndoVis) challenge.
2 papers · 0 benchmarks
FinBench is a benchmark for evaluating the performance of machine learning models with both tabular data inputs and profile text inputs.
2 papers · 0 benchmarks
Fingerprint Dataset (Neural Audio Fingerprint Dataset)
This dataset includes all music sources, background noises and impulse-reponses (IR) samples and conversation speech that have been used in the work "Neural Audio Fingerprint for High-specific Audio Retrieval based on Contrastive Learning"…
2 papers · 0 benchmarks
FormulaNet FormulaNet is a new large-scale Mathematical Formula Detection dataset.
2 papers · 0 benchmarks
FracAtlas (A Dataset for Fracture Classification, Localization and Segmentation of Musculoskeletal Radiographs)
FractureAtlas is a musculoskeletal bone fracture dataset with annotations for deep learning tasks like classification, localization, and segmentation.
2 papers · 0 benchmarks
This dataset is dialog dataset collected in a Wizard-of-Oz fashion.
2 papers · 0 benchmarks
GATITOS (Google's Additional Translations Into Tail-languages: Often Short)
The GATITOS (Google's Additional Translations Into Tail-languages: Often Short) dataset is a high-quality, multi-way parallel dataset of tokens and short phrases, intended for training and improving machine translation models.
2 papers · 0 benchmarks
GEOBench-VLM, a comprehensive benchmark specifically designed to evaluate VLMs on geospatial tasks, including scene understanding, object counting, localization, fine-grained categorization, and temporal analysis.
2 papers · 0 benchmarks
GIRT-Data (GitHub Issue Report Template Dataset)
GIRT-Data is the first and largest dataset of issue report templates (IRTs) in both YAML and Markdown format.
2 papers · 0 benchmarks
GIS (Github Issue Similarity)
This dataset can be used for semantic textual similarity tasks.
2 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.