Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 79 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 3745–3792 of 3,998
Cancer genomics and precision oncology: The TCGA Research Network started in 2005 has profiled and analyzed a large number of human tumors to discover molecular aberrations at the DNA, RNA, protein, and epigenetic levels and thereby…
0 papers · 0 benchmarks
The landmark Cancer Genomics Program launched in 2006 has contributed immensely to the awareness of the importance of cancer genomics in our understanding of cancer over the past decade and has begun to change the way the disease is…
0 papers · 0 benchmarks
The DeepSpeak dataset contains over 43 hours of real and deepfake footage of people talking and gesturing in front of their webcams.
0 papers · 0 benchmarks
DigiLeTs (Digit- and Letter Trajectories)
A dataset with 23 870 digital trajectories (i.e.
0 papers · 0 benchmarks
This dataset is based on FB15k237 and a pre-trained language-model-based KGE.
0 papers · 0 benchmarks
This dataset contains electroencephalogram (EEG) signals, event-related potentials (ERP), and demographic attributes aimed at the early identification of schizophrenia.
0 papers · 0 benchmarks
ESP dataset (Evaluation for Styled Prompt dataset) is a new benchmark for zero-shot domain-conditional caption generation.
0 papers · 0 benchmarks
The English-Pashto Language Dataset (EPLD) is a comprehensive resource aimed to provide linguistic insights into the Pashto language.
0 papers · 0 benchmarks
A dataset specifically tailored to the biotech news sector, aiming to transcend the limitations of existing benchmarks.
0 papers · 0 benchmarks
This is a machine-learning-ready glaucoma dataset using a balanced subset of standardized fundus images from the Rotterdam EyePACS AIROGS train set.
0 papers · 0 benchmarks
This is an improved machine-learning-ready glaucoma dataset using a balanced subset of standardized fundus images from the Rotterdam EyePACS AIROGS [1] set.
0 papers · 0 benchmarks
FALLMUD (FAscicle Lower Leg Muscle Ultrasound Dataset)
FAscicle Lower Leg Muscle Ultrasound Dataset is a dataset composed of 812 ultrasound images of lower leg muscles to analyze muscle weaknesses and prevent injuries.
0 papers · 0 benchmarks
FHRMA is an open-source project for Fetal Heart Rate Morphological Analysis containing Matlab source code and datasets.
0 papers · 0 benchmarks
FSI (Fluid-Solid interaction)
Data Set Structure Fluid Structure Interaction(NS +Elastic wave) The TFfsi2results folder contains simulation data organized by various parameters (mu, x1, x2) where mu determines the viscosity and x1 and x2 are the parameters of the inlet…
0 papers · 0 benchmarks
The free Face dataset made for students and teachers.
0 papers · 0 benchmarks
We introduce an annotated dataset of five thousand human labeled pareidolic face images, called Faces in Things''.
0 papers · 0 benchmarks
The Fields2Benhmark dataset is a collection of 350 agricultural fields in vector format manually selected to test agricultural coverage path planning algorithms.
0 papers · 0 benchmarks
FinArg (Financial Argument Mining)
With the goal of reasoning on the financial textual data, we present a novel dataset for annotating arguments, their components, and relations in the transcripts of earnings conference calls (ECCs).
0 papers · 0 benchmarks
FluencyBank is a shared database for the study of fluency development.
0 papers · 0 benchmarks
We introduce FortisAVQA, a dataset designed to assess the robustness of AVQA models.
0 papers · 0 benchmarks
This record contains the saddle search output logs for Sella and EON (dimer, with and without GPR acceleration).
0 papers · 0 benchmarks
Gap Pattern Detection (Gap Pattern (Gap Up and Gap Down) Detection in Candlestick Trading Charts for Technical Analysis)
1.
0 papers · 0 benchmarks
GenAI-Bench benchmark consists of 1,600 challenging real-world text prompts sourced from professional designers.
0 papers · 0 benchmarks
We've made available several genome-wide datasets, which can be used for training microRNA (miRNA) classifiers.
0 papers · 0 benchmarks
KOKLU Murat (a), UNLERSEN M.
0 papers · 0 benchmarks
HEADSET (HEADSET: Human Emotion Awareness under Partial Occlusions Multimodal DataSET)
The volumetric representation of human interactions is one of the fundamental domains in the development of immersive media productions and telecommunication applications.
0 papers · 0 benchmarks
The medaka (Oryzias latipes) and the zebrafish (Danio rerio) are used as a model organism for a variety of subjects in biomedical research.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.