Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 49 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 2305–2352 of 3,998

A real-world low-light camera motion blur dataset for evaluating deblurring radiance fields methods.
1 paper · 0 benchmarks
The first large-scale dataset for training and evaluating novel-view synthesis from blurred images.
1 paper · 0 benchmarks
Bone Cement Removal with Audio-Monitoring (Bone Cement Removal with Audio-Monitoring and Erosion Depth)
This dataset comprises extensive multi-modal data related to the experimental study of ultrasonically excited pulsating fluid jets used for bone cement removal.
1 paper · 0 benchmarks
Boreal Forest Fire (Boreal Forest Fire: UAV-collected Wildfire Detection and Smoke Segmentation Dataset)
This dataset consists of annotated images and videos of smoke resulting from prescribed burning events in Finnish boreal forests.
1 paper · 0 benchmarks
RGB-D instance segmentation box dataset.
1 paper · 2 benchmarks
BraTS PEDs 2023 (The Brain Tumor Segmentation (BraTS) Challenge 2023: Focus on Pediatrics (CBTN-CONNECT-DIPGR-ASNR-MICCAI BraTS-PEDs))
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
BraTS-Africa (Brain Tumor Segmentation (BraTS) Challenge: Sub Saharan Africa)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 1 benchmark
BraTs Peds 2024 (The Brain Tumor Segmentation in Pediatrics (BraTS-PEDs) Challenge (CBTN-CONNECT-DIPGR-ASNR-MICCAI BraTS-PEDs) 2024)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 1 benchmark
Brain2Text (Data for: A high-performance speech neuroprosthesis)
From the dataset paper: Brain-computer interfaces (BCIs) can restore communication to people who have lost the ability to move or speak.
1 paper · 0 benchmarks
Bramble flower image dataset (BRAMBLE FLOWER DETECTION AND CLASSIFICATION DATASET FOR PRECISION POLLINATION)
This dataset contains both the artificial and real flower images of bramble flowers.
1 paper · 0 benchmarks
See https://www.kaggle.com/datasets/olistbr/brazilian-ecommerce .
1 paper · 0 benchmarks
BreastDICOM4 ([MIMBCD-UI] UTA4: Medical Imaging DICOM Files Dataset)
Several datasets are fostering innovation in higher-level functions for everyone, everywhere.
1 paper · 1 benchmark
BreastRates4 ([MIMBCD-UI] UTA4: Rates Dataset)
Several datasets are fostering innovation in higher-level functions for everyone, everywhere.
1 paper · 0 benchmarks
Dataset of 5,591 labeled issue tickets.
1 paper · 0 benchmarks
The original paper contains a high-level explanation of the dataset characteristics, and potential use cases of the dataset.
1 paper · 0 benchmarks
BurnMD (A Fire Projection and Mitigation Modeling Dataset)
A dataset composed of 308 medium sized fires from the years 2018-2021, complete with both time series airborne based inference and ground operational estimation of fire extent, and operational mitigation data such as control line…
1 paper · 0 benchmarks
This dataset contains the bus trajectory dataset collected by 6 volunteers who were asked to travel across the sub-urban city of Durgapur, India, on intra-city buses (route name: 54 Feet).
1 paper · 0 benchmarks
Business license datasets and source code for named entity recognition.
1 paper · 0 benchmarks
This is the C++ dataset used in the TASTY research paper which was published at the ICLR DL4Code (Deep Learning for Code) workshop.
1 paper · 0 benchmarks
The feature files are named with the youtube IDs.
1 paper · 0 benchmarks
📚 CADBench CADBench is a comprehensive benchmark to evaluate the ability of LLMs to generate CAD scripts.
1 paper · 0 benchmarks
CAESAR-Radi (CAESAR-Radi: SAR-Ship-Dataset)
This dataset labeled by SAR experts was created using 102 Chinese Gaofen-3 images and 108 Sentinel-1 images.
1 paper · 0 benchmarks
CANDOR Corpus (CANDOR = Conversation: A Naturalistic Dataset of Online Recordings)
The CANDOR corpus is a large, novel, multimodal corpus of 1,656 recorded conversations in spoken English.
1 paper · 0 benchmarks
This dataset contains synthetic images extracted from the CARLA simulator along with rich information extracted from the deferred rendering pipeline of Unreal Engine 4.
1 paper · 0 benchmarks
CASA (Clinical annotations for automatic Stuttering Assessment)
Annotation Guidlines Three speech and language pathologists, with experience ranging from 2 to 40 years, independently annotated and analyzed the audiovisual samples sourced from the Fluencybank Adults Who Stutter(AWS) dataset.
1 paper · 0 benchmarks
CASTLE Benchmark (CASTLE Benchmark C@250)
The CASTLE Benchmark is a comprehensive dataset and a scoring method for evaluating single or combinations of static analyzers with a focus on security.
1 paper · 0 benchmarks
CAShift (Cloud Attack & Normality Shift Dataset)
CAShift is the first multiple normality shift-aware Log-Based Anomaly Detection (LAD) dataset specifically designed for cloud systems, which considers different software roles in cloud systems and attack behavior among cloud components.
1 paper · 0 benchmarks
CAT is a specialized dataset for co-saliency detection.
1 paper · 0 benchmarks
CAsT-answerability dataset contains binary answerability labels on three levels: sentence, passage, and ranking.
1 paper · 0 benchmarks
CCIHP (Characterized Crowd Instance-level Human Parsing)
CCIHP dataset is devoted to fine-grained description of people in the wild with localized & characterized semantic attributes.
1 paper · 0 benchmarks
CCPT (Conceptual Combination with Property Type)
CCPT is a dataset containing 12.3K triplets of noun phrases, properties, and property types for conceptual combination.
1 paper · 0 benchmarks
CCTSDB-AUG (From CCSPNet-Joint,IJCNN 2024)
The CSUST Chinese Traffic Sign Detection Benchmark (CCTSDB) is an existing dataset for traffic sign detection.
1 paper · 1 benchmark
CD-HARD comprises 102 images featuring vehicles with oblique license plates sourced from the Cars dataset.
1 paper · 0 benchmarks
CECW (Colorful Extended Cleanup World)
The CECW dataset is a color-extended version of the Cleanup World (CW) borrowed from the mobile-manipulation robot domain.
1 paper · 0 benchmarks
CEREC (Corpus for Entity Resolution in Email Conversations)
CEREC is a large scale corpus for entity resolution in email conversations.
1 paper · 0 benchmarks
CGHD1152 (Circuit Graph Hand Drawn 1152)
- 1152 Images - 144 Circuits - 12 Drafter - 48,563 Object (Symbol, Structural, Text) Annotations
1 paper · 0 benchmarks
An image sequence dataset of growing snowflakes in HDF5 format.
1 paper · 0 benchmarks
CHAMP (Concept and Hint-Annotated Math Problems)
The Concept and Hint-Annotated Math Problems (CHAMP) consists of high school math competition problems, annotated with concepts, or general math facts, and hints, or problem-specific tricks.
1 paper · 0 benchmarks
CI-ToD is a dataset for Consistency Identification in Task-oriented Dialog system.
1 paper · 0 benchmarks
CIP (Complete Inertial Pose)
The CIP dataset is composed of 2 subsets, containing low-cost (MPU9250) and high-end (MTwAwinda) Magnetic, Angular Rate, and Gravity (MARG) sensor data respectively.
1 paper · 0 benchmarks
Description This repository includes the experiment results, source code, and test data for Three Cs risk inference, using the CIRO (COVID-19 Infection Risk Ontology) and HermiT.
1 paper · 0 benchmarks
CKBP v2 is a new CSKB Population benchmark, which addresses the two mentioned problems by using experts instead of crowd-sourced annotation and by adding diversified adversarial samples to make the evaluation set more representative.
1 paper · 0 benchmarks
CLCXray (Cutters and Liquid Containers X-ray Dataset)
The CLCXray dataset contains 9,565 X-ray images, in which 4,543 X-ray images (real data) are obtained from the real subway scene and 5,022 X-ray images (simulated data) are scanned from manually designed baggage.
1 paper · 1 benchmark
CLEVR-MRT (CLEVR: Mental Rotation Tests)
CLEVR Mental Rotation Tests (CLEVR-MRT) is a new version of the CLEVR dataset.
1 paper · 0 benchmarks
CLUES is a benchmark for Classifier Learning Using natural language ExplanationS, consisting of a range of classification tasks over structured data along with natural language supervision in the form of explanations.
1 paper · 0 benchmarks
Contains a dataset of 241 Chinese dishes with 191,811 images.
1 paper · 0 benchmarks
CNFOOD-241 Contains a dataset of 241 Chinese dishes with 191,811 images.
1 paper · 1 benchmark
COAT (CommonSense Object Affordance Task)
Useful for checking the physical reasoning capabilities in household agents.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.