Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 53 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 2497–2544 of 3,998
This is the dataset used for classifying Gene-Disease relationship types from sentences.
1 paper · 1 benchmark
This dataset contains recordings of 32 sound producing insect species with a total 335 files and a length of 57 minutes.
1 paper · 0 benchmarks
The laparoscopic surgery dataset is associated with our International Journal of Computer Assisted Radiology and Surgery (IJCARS) publication titled “DeSmoke-LAP: Improved Unpaired Image-to-Image Translation for Desmoking in Laparoscopic…
1 paper · 0 benchmarks
DeVAn (Dense Video Annotation for Video-Language Models)
DeVAn is a multi-modal dataset containing 8.5K video clips carefully selected from previously published YouTube-based video datasets (YouTube-8M and YT-Temporal-1B) that integrate visual and auditory information.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This dataset inclue multi-spectral acquisition of vegetation for the conception of new DeepIndices.
1 paper · 1 benchmark
Release for uploading scripts and data to Zenodo Deep Neural Network Training Script incl.
1 paper · 0 benchmarks
The dataset contains two Pareto-fronts: - The Pareto-front for the 2-objective problem - The Pareto-front for the 3-objective problem Each Pareto-front contains a set of points, with coordinates given by their objectives.
1 paper · 0 benchmarks
There are 537 RGB jpg images of cracks and corresponding png binary segmentation masks of crack: a training set with 300 images and a testing set with 237 images.
1 paper · 0 benchmarks
DeepGraviLens is a data set of simulated gravitational lenses consisting of images associated with brightness variation time series.
1 paper · 0 benchmarks
DeepParliament is a legal domain Benchmark Dataset that gathers bill documents and metadata and performs various bill status classification tasks.
1 paper · 0 benchmarks
Demonstration video of the Stickbug Robot
1 paper · 0 benchmarks
Corpus for argument mining in legal documents, composed of 40 decisions of the Court of Justice of the European Union on matters of fiscal state aid
1 paper · 0 benchmarks
Provide: 10 pickle files 8 pickle files are used to generate depth maps 2 pickle files are data of fronto parallel texture with their ground truth depth Each pickle file contains the parameter of the camera system (aperture size, optical…
1 paper · 0 benchmarks
A dataset of 100K synthetic images of skin lesions, ground-truth (GT) segmentations of lesions and healthy skin, GT segmentations of seven body parts (head, torso, hips, legs, feet, arms and hands), and GT binary masks of non-skin regions…
1 paper · 0 benchmarks
A curated dataset of 221 question-answer-rationale triples capturing visualization design decisions and the reasoning behind them, derived from real-world student-authored narratives.
1 paper · 0 benchmarks
The code that created this dataset can be seen in https://github.com/nitzanfarhi/SecurityPatchDetection and can be reproduced by running: console python datacollection\createdataset.py --all -o datacollection\data Notice that this dataset…
1 paper · 0 benchmarks
DialogCC is a large-scale multi-modal dialogue dataset, which covers diverse real-world topics and various images per dialogue.
1 paper · 0 benchmarks
This dataset curates quantitative transparency disclosures about the online sexual exploitation of minors.
1 paper · 0 benchmarks
Dataset Summary The DiscoEval is an English-language Benchmark that contains a test suite of 7 tasks to evaluate whether sentence representations include semantic information relevant to discourse processing.
1 paper · 0 benchmarks
Project: Discrete-Time Modeling of Interturn Short Circuits in Interior PMSMs Authors: Lukas Zezula, Matus Kozovsky, Ludek Buchta and Petr Blaha (corresponding author: Lukas Zezula, e-mail: lukas.zezula@ceitec.vutbr.cz) Affiliation: CEITEC…
1 paper · 0 benchmarks
Dissonance Twitter Dataset is a dataset collected from annotating tweets for dissonance.
1 paper · 0 benchmarks
This dataset is named as the DistNLI dataset, which is a synthesized benchmark aiming to probe neural network models from the aspect of conjunctions on distributivity in NLI task in American English.
1 paper · 0 benchmarks
DnR-nonverbal is a dataset for cinematic audio source separation (CASS) based on Divide and Remaster (DnR) dataset.
1 paper · 0 benchmarks
This dataset consisting 500 set of caption, table and coresponding paper page, processed from DocBank.
1 paper · 0 benchmarks
We manually annotate 800 sentences from 80 documents in two domains (Healthcare and Transportation) to form a DocOIE dataset for evaluation.
1 paper · 2 benchmarks
DocRED-FE (DocRED with Fine-Grained Entity Type)
DocRED-FE is the DocRED with Fine-Grained Entity Type
1 paper · 0 benchmarks
The DocRED Information Extraction (DocRED-IE) dataset extends the DocRED dataset for the Document-level Closed Information Extraction (DocIE) task.
1 paper · 6 benchmarks
The dataset is composed of 95 unique document texts spanning the period 2005-2022.
1 paper · 0 benchmarks
The MLCommons Dollar Street Dataset is a collection of images of everyday household items from homes around the world that visually captures socioeconomic diversity of traditionally underrepresented populations.
1 paper · 0 benchmarks
Abstract Forecasting methods from averaging regression analysis lines, reversed and direct lines.
1 paper · 0 benchmarks
The Drag100 dataset is introduced in the paper "GoodDrag: Towards Good Practices for Drag Editing with Diffusion Models"¹.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
About the Dataset: 4 classes of drinking waste: Aluminium Cans, Glass bottles, PET (plastic) bottles and HDPE (plastic) Milk bottles.
1 paper · 1 benchmark
A synthetic dataset including driving under adverse weather conditions | Autonomous Driving
1 paper · 0 benchmarks
人群计数旨在识别物体的数量,在智能交通、城市管理和安全监控中发挥着重要作用。由于比例变化、照明变化、遮挡和较差的成像条件,尤其是在夜间和雾霾条件下,人群计数的任务非常具有挑战性。 在本文中,我们提出了一个基于无人机的 RGB-Thermal 人群计数数据集 (DroneRGBT),该数据集由 3600…
1 paper · 1 benchmark
The data used for all results in this paper can be found here.
1 paper · 0 benchmarks
Dubbing Test Set consists of two subsets extracted from the En→De test set of COVOST-2, a large-scale multilingual speech translation corpus based on Common Voice.
1 paper · 0 benchmarks
As Large Language Models (LLMs) increasingly participate in human-AI interactions, evaluating their Theory of Mind (ToM) capabilities - particularly their ability to track dynamic mental states - becomes crucial.
1 paper · 0 benchmarks
This dataset is based on WN18RR and a pre-trained language-model-based KGE.
1 paper · 0 benchmarks
E2E Refined is a dataset for sentence classification.
1 paper · 0 benchmarks
EA-HAS-Bench (Energy-Aware Hyperparameter and Architecture Search Benchmark)
We present the first large-scale energy-aware benchmark that allows studying AutoML methods to achieve better trade-offs between performance and search energy consumption, named EA-HAS-Bench.
1 paper · 0 benchmarks
EBHI-Seg is a dataset containing 5,170 images of six types of tumor differentiation stages and the corresponding ground truth images.
1 paper · 0 benchmarks
ECTF (Early COVID-19 Twitter Fake news)
ECTF is a dataset for Twitter fake news detection in the Covid-19 domain.
1 paper · 0 benchmarks
This dataset is built from 10-Q documents (Quarterly Reports) of publicly listed companies on the SEC.
1 paper · 0 benchmarks
All data is from one continuous EEG measurement with the Emotiv EEG Neuroheadset.
1 paper · 0 benchmarks
EGC-FPHFS (Early Gastric Cancer Data from First People's Hospital of Foshan)
High-resolution early gastric cancer (EGC) detection and analysis: Patient Data:Datasets often include images from patients diagnosed with gastric cancer, specifically distinguishing between early gastric cancer (EGC) and Non -pathogenic…
1 paper · 1 benchmark
This dataset contains transcriptions of the electric guitar performance of 240 tablatures, rendered with different tones.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.