Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 42 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 1969–2016 of 3,998

New Plant Diseases Dataset (Image dataset containing different healthy and unhealthy crop leaves.)
This dataset is recreated using offline augmentation from the original dataset.
2 papers · 2 benchmarks
News Interactions on Globo.com (News Portal User Interactions by Globo.com - A large dataset for news recommendations offline evaluation and analytics)
Context This large dataset with users interactions logs (page views) from a news portal was kindly provided by [Globo.com][1], the most popular news portal in Brazil, for reproducibility of the experiments with CHAMELEON - a…
2 papers · 0 benchmarks
NoW (Noise of Web)
Noise of Web (NoW) is a challenging noisy correspondence learning (NCL) benchmark for robust image-text matching/retrieval models.
2 papers · 0 benchmarks
Abstract The Norwegian Endurance Athlete ECG Database contains 12-lead ECG recordings from 28 elite athletes from various sports in Norway.
2 papers · 0 benchmarks
Scene-focused, multi-modal, episodic data of the images and symbolic world-states seen by an agent completing a pogo-stick assembly task within a video game world.
2 papers · 0 benchmarks
OAD dataset (The Online Action Detection Dataset)
The Online Action Detection Dataset (OAD) was captured using the Kinect V2 sensor, which collects color images, depth images and human skeleton joints synchronously.
2 papers · 1 benchmark
OADAT (OADAT: Experimental and Synthetic Clinical Optoacoustic Data for Standardized Image Processing)
An experimental and synthetic (simulated) OA raw signals and reconstructed image domain datasets rendered with different experimental parameters and tomographic acquisition geometries.
2 papers · 0 benchmarks
OCR-IDL (OCR Annotations for Industry Document Library Dataset)
The OCR-IDL dataset comprises the OCR annotations for a subset of 26M pages of the large-scale IDL document library.
2 papers · 0 benchmarks
A dataset containing the results of a MUSHRA listening test conducted with expert listeners from 2 international laboratories.
2 papers · 1 benchmark
ORCAS-I (Queries Annotated with Intent using Weak Supervision)
A labelled version of the ORCAS click-based dataset of Web queries, which provides 18 million connections to 10 million distinct queries.
2 papers · 1 benchmark
Occluded COCO is automatically generated subset of COCO val dataset, collecting partially occluded objects for a large variety of categories in real images in a scalable manner, where target object is partially occluded but the…
2 papers · 1 benchmark
Ocean Drifters (Madagascar Ocean Drifters)
From Schaub, Michael T., et al.
2 papers · 0 benchmarks
This is a transnational data set which contains all the transactions occurring between 01/12/2010 and 09/12/2011 for a UK-based and registered non-store online retail.
2 papers · 0 benchmarks
Only Time Will Tell (Time-respecting and time-ignoring horizon of code review network at Microsoft)
Simulation results of time-respecting and time-ignoring horizon of code review network at Microsoft as JSON.
2 papers · 0 benchmarks
OpenAsp Dataset OpenAsp is an Open Aspect-based Multi-Document Summarization dataset derived from DUC and MultiNews summarization datasets.
2 papers · 0 benchmarks
(L)ifel(O)ng (R)obotic V(IS)ion (OpenLORIS) - Object Recognition Dataset (OpenLORIS-Object) is designed for accelerating the lifelong/continual/incremental learning research and application,currently focusing on improving the continuous…
2 papers · 0 benchmarks
OpenSLR (Open Speech and Language Resources)
OpenSLR is a repository of open speech and language resources, including large-scale transcribed audio corpora and related software.
2 papers · 1 benchmark
OpenViDial 2.0 is a larger-scale open-domain multi-modal dialogue dataset compared to the previous version OpenViDial 1.0.
2 papers · 1 benchmark
PAD Dataset (Pose-agnostic/Multi-pose Anomaly Detection Dataset)
Multi-pose Anomaly Detection (MAD) dataset, which represents the first attempt to evaluate the performance of pose-agnostic anomaly detection.
2 papers · 1 benchmark
Enables research on early detection of sexual predators in chats (eSPD).
2 papers · 0 benchmarks
Appearance-based gaze estimation systems have shown great progress recently, yet the performance of these techniques depend on the datasets used for training.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
The dataset contains 45 documents containing narrative description of business process and their annotations.
2 papers · 0 benchmarks
PETCI (PETCI: A Parallel English Translation Dataset of Chinese Idioms)
PETCI is a Parallel English Translation dataset of Chinese Idioms, collected from an idiom dictionary and Google and DeepL translation.
2 papers · 0 benchmarks
The PKU dataset has almost 4,000 images categorized into five groups (G1-G5) that show different situations.
2 papers · 0 benchmarks
The dataset used in the experiments on the paper "Modeling citation worthiness by using attention‑based bidirectional long short‑term memory networks and interpretable models" There are one million sentences in total, and further splitted…
2 papers · 0 benchmarks
POINTREC is a test collection for point of interest (POI) recommendation, comprising of (i) a set of information needs, (ii) a dataset of POIs, and (iii) graded relevance assessments for information need and POI pairs.
2 papers · 0 benchmarks
PPG Dalia (PPG Field Study Dataset)
PPG-DaLiA is a publicly available dataset for PPG-based heart rate estimation.
2 papers · 0 benchmarks
Automated leaf segmentation is a challenging area in computer vision.
2 papers · 0 benchmarks
Dataset Card for The Cancer Genome Atlas (TCGA) Multimodal Dataset The Cancer Genome Atlas (TCGA) Multimodal Dataset is a comprehensive collection of clinical data, pathology reports, molecular, and slide images for cancer patients.
2 papers · 0 benchmarks
Pansharpening Datasets from WorldView 2, WorldView 3, QuickBird, Gaofen 2 sensors.
2 papers · 4 benchmarks
Pano3D is a new benchmark for depth estimation from spherical panoramas.
2 papers · 0 benchmarks
Paper2Fig100k is a dataset with over 100k images of figures and texts from research papers.
2 papers · 0 benchmarks
ParaMAWPS (Paraphrased Math Word Problem Solving Repository)
This repository contains the code, data, and models of the paper titled "Math Word Problem Solving by Generating Linguistic Variants of Problem Statements" published in the Proceedings of the 61st Annual Meeting of the Association for…
2 papers · 1 benchmark
To take advantage of the ever-increasing amount of structural data now available, we also trained Paragraph on a larger dataset.
2 papers · 1 benchmark
Perseus is a dataset for Cross-Lingual Summarization (CLS) which collects about 94K Chinese scientific documents paired with English summaries.
2 papers · 0 benchmarks
PersonPath22 is a large-scale multi-person tracking dataset containing 236 videos captured mostly from static-mounted cameras, collected from sources where we were given the rights to redistribute the content and participants have given…
2 papers · 1 benchmark
An annotated dataset of 38,800 phishing and benign websites.
2 papers · 0 benchmarks
Introduction The 2016 PhysioNet/CinC Challenge aims to encourage the development of algorithms to classify heart sound recordings collected from a variety of clinical or nonclinical (such as in-home visits) environments.
2 papers · 0 benchmarks
145k natural language and PDDL problem pairs from the Blocks World, Gripper, and Floor Tile domains.
2 papers · 0 benchmarks
PoKi is a corpus of 61,330 poems written by children from grades 1 to 12.
2 papers · 0 benchmarks
PoliteRewrite (the politerewrite dataset)
https://huggingface.co/datasets/jdustinwind/Polite
2 papers · 0 benchmarks
PolyDensity (Polymer Density)
The PolyDensity is collected from Polyinfo.
2 papers · 0 benchmarks
PropSegmEnt is a corpus of over 35K propositions annotated by expert human raters.
2 papers · 0 benchmarks
Psychometric NLP is a corpus for psychometric natural language processing (NLP) related to important dimensions such as trust, anxiety, numeracy, and literacy, in the health domain.
2 papers · 0 benchmarks
A collection of 385,705 scientific abstracts about Cognitive Control and their GPT-3 embeddings.
2 papers · 1 benchmark
Datasets of QCD jets used for studying unfolding in OmniFold: A Method to Simultaneously Unfold All Observables.
2 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.