Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 82 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 3889–3936 of 3,998

Dataset: RGB-D Images for Real-World and Synthetic Object Scenes This dataset consists of both real-world and synthetic RGB-D images, designed for object detection, classification, and segmentation tasks, particularly for primitive shape…
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
RTI Rwanda Drone Crop Types (Drone Imagery Classification Training Dataset for Crop Types in Rwanda)
RTI International (RTI) generated 2,611 labeled point locations representing 19 different land cover types, clustered in 5 distinct agroecological zones within Rwanda.
0 papers · 0 benchmarks
A review on raw subjective scores and data manipulation for before and after refining Mean opinion Scores
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
This dataset contains a curated sample of 20,000 English-language user reviews sourced exclusively from Trustpilot.com.
0 papers · 0 benchmarks
Citation Request: See the articles for more detailed information on the data.
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
SEN (Sentiment analysis of Entities in News headlines)
SEN is a novel publicly available human-labelled dataset for training and testing machine learning algorithms for the problem of entity level sentiment analysis of political news headlines.
0 papers · 0 benchmarks
Dataset Card for SENTINEL: Mitigating Object Hallucinations via Sentence-Level Early Intervention For the details of this dataset, please refer to the documentation of the GitHub repo.
0 papers · 0 benchmarks
Facial landmark detection is a cornerstone in many facial analysis tasks such as face recognition, drowsiness detection, and facial expression recognition.
0 papers · 0 benchmarks
SICS-155 (Phase Recognition in Small Incision Cataract Surgery Videos)
Cataract is the leading cause of blindness worldwide, most affecting life in low- and middle-income countries (LMICs).
0 papers · 0 benchmarks
SILD (Survey Item Linking Dataset)
This dataset contains a collection of texts from publications from a broad range of social science domains (e.g., economics, politics, psychology, etc.).
0 papers · 0 benchmarks
SMCOVID19-CT (Contact Tracing Data (from Italian SM-COVID-19 App))
We present a real data analysis of a CT experiment that was conducted in Italy for 8 months and involved more than 100,000 CT app users.
0 papers · 0 benchmarks
SMDG (Standardized Multi-Channel Dataset for Glaucoma)
Standardized Multi-Channel Dataset for Glaucoma (SMDG-19) is a collection and standardization of 19 public datasets, comprised of full-fundus glaucoma images, associated image metadata like, optic disc segmentation, optic cup segmentation,…
0 papers · 0 benchmarks
Grounding Scientific Entity References in STEM Scholarly Content to Authoritative Encyclopedic and Lexicographic Sources The STEM ECR v1.0 dataset has been developed to provide a benchmark for the evaluation of scientific entity…
0 papers · 0 benchmarks
Saarbruecken Voice Database contains voice and EGG recordings of patients diagnosed with voice disorder, as well as healthy persons.
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
This dataset contains images and annotations for scene text detection and recognition.
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
SourceData-NLP (The SourceData-NLP dataset: integrating curation into scientific publishing for training large language models)
Introduction: The scientific publishing landscape is expanding rapidly, creating challenges for researchers to stay up-to-date with the evolution of the literature.
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.