Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 82 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 3889–3936 of 3,998
Dataset: RGB-D Images for Real-World and Synthetic Object Scenes This dataset consists of both real-world and synthetic RGB-D images, designed for object detection, classification, and segmentation tasks, particularly for primitive shape…
0 papers · 0 benchmarks
RTI International (RTI) generated 2,611 labeled point locations representing 19 different land cover types, clustered in 5 distinct agroecological zones within Rwanda.
0 papers · 0 benchmarks
A review on raw subjective scores and data manipulation for before and after refining Mean opinion Scores
0 papers · 0 benchmarks
This dataset contains a curated sample of 20,000 English-language user reviews sourced exclusively from Trustpilot.com.
0 papers · 0 benchmarks
Citation Request: See the articles for more detailed information on the data.
0 papers · 0 benchmarks
SEN (Sentiment analysis of Entities in News headlines)
SEN is a novel publicly available human-labelled dataset for training and testing machine learning algorithms for the problem of entity level sentiment analysis of political news headlines.
0 papers · 0 benchmarks
Dataset Card for SENTINEL: Mitigating Object Hallucinations via Sentence-Level Early Intervention For the details of this dataset, please refer to the documentation of the GitHub repo.
0 papers · 0 benchmarks
Facial landmark detection is a cornerstone in many facial analysis tasks such as face recognition, drowsiness detection, and facial expression recognition.
0 papers · 0 benchmarks
SICS-155 (Phase Recognition in Small Incision Cataract Surgery Videos)
Cataract is the leading cause of blindness worldwide, most affecting life in low- and middle-income countries (LMICs).
0 papers · 0 benchmarks
SILD (Survey Item Linking Dataset)
This dataset contains a collection of texts from publications from a broad range of social science domains (e.g., economics, politics, psychology, etc.).
0 papers · 0 benchmarks
SMCOVID19-CT (Contact Tracing Data (from Italian SM-COVID-19 App))
We present a real data analysis of a CT experiment that was conducted in Italy for 8 months and involved more than 100,000 CT app users.
0 papers · 0 benchmarks
SMDG (Standardized Multi-Channel Dataset for Glaucoma)
Standardized Multi-Channel Dataset for Glaucoma (SMDG-19) is a collection and standardization of 19 public datasets, comprised of full-fundus glaucoma images, associated image metadata like, optic disc segmentation, optic cup segmentation,…
0 papers · 0 benchmarks
Grounding Scientific Entity References in STEM Scholarly Content to Authoritative Encyclopedic and Lexicographic Sources The STEM ECR v1.0 dataset has been developed to provide a benchmark for the evaluation of scientific entity…
0 papers · 0 benchmarks
Saarbruecken Voice Database contains voice and EGG recordings of patients diagnosed with voice disorder, as well as healthy persons.
0 papers · 0 benchmarks
This dataset contains images and annotations for scene text detection and recognition.
0 papers · 0 benchmarks
SourceData-NLP (The SourceData-NLP dataset: integrating curation into scientific publishing for training large language models)
Introduction: The scientific publishing landscape is expanding rapidly, creating challenges for researchers to stay up-to-date with the evolution of the literature.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.