Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 35 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 1633–1680 of 3,998

PCD (Poem Comprehensive Dataset)
The Arabic dataset is scraped mainly from الموسوعة الشعرية and الديوان.
3 papers · 3 benchmarks
PDE dataset (Parametric Partial Differential Equation dataset)
Contains data of parametric PDEs - Burgers' equation - Darcy's flow - Navier-Stokes equation
3 papers · 0 benchmarks
PDFVQA: A New Dataset for Real-World VQA on PDF Documents
3 papers · 0 benchmarks
PETRAW (PEg TRAnsfer Workflow recognition by different modalities)
PETRAW data set was composed of 150 sequences of peg transfer training sessions.
3 papers · 6 benchmarks
PNT (Parsing Time Normalizations)
The Parsing Time Normalizations (PNT) corpus in SCATE format allows the representation of a wider variety of time expressions than previous approaches.
3 papers · 1 benchmark
POPGym (Partially Observable Process Gym)
POPGym is designed to benchmark memory in deep reinforcement learning.
3 papers · 0 benchmarks
This dataset contains both infeasible and feasible data points as described in PRIME.
3 papers · 0 benchmarks
PWC Leaderboards (Papers with Code Leaderboards)
The Papers with Code Leaderboards dataset is a collection of over 5,000 results capturing performance of machine learning models.
3 papers · 1 benchmark
PWDB (Pulse Wave Database)
Overview This database of simulated arterial pulse waves is designed to be representative of a sample of pulse waves measured from healthy adults.
3 papers · 0 benchmarks
This paper presents a benchmark data set for condition monitoring of rolling bearings in combination with an extensive description of the corresponding bearing damage, the data set generation by experiments and results of datadriven…
3 papers · 0 benchmarks
Modeling what makes an advertisement persuasive, i.e., eliciting the desired response from consumer, is critical to the study of propaganda, social psychology, and marketing.
3 papers · 0 benchmarks
PoPArt (Poses of People in Art: A Data Set for Human Pose Estimation in Digital Art History)
Throughout the history of art, the pose—as the holistic abstraction of the human body's expression—has proven to be a constant in numerous studies.
3 papers · 1 benchmark
Polyps in the colon are widely known cancer precursors identified by colonoscopy.
3 papers · 1 benchmark
PubChemQA consists of molecules and their corresponding textual descriptions from PubChem.
3 papers · 1 benchmark
PubMedCite is a domain-specific dataset with about 192K biomedical scientific papers and a large citation graph preserving 917K citation relationships between them.
3 papers · 0 benchmarks
QALD-9-Plus Dataset Description QALD-9-Plus is the dataset for Knowledge Graph Question Answering (KGQA) based on well-known QALD-9.
3 papers · 1 benchmark
RETWEET is a dataset of tweets and overall predominant sentiment of their replies.
3 papers · 2 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 0 benchmarks
ROOR is a reading order prediction (ROP) benchmark which annotates layout reading order as ordering relations.
3 papers · 1 benchmark
The RailEye3D dataset, a collection of train-platform scenarios for applications targeting passenger safety and automation of train dispatching, consists of 10 image sequences captured at 6 railway stations in Austria.
3 papers · 0 benchmarks
The RareDis corpus contains more than 5,000 rare diseases and almost 6,000 clinical manifestations are annotated.
3 papers · 0 benchmarks
ResQ (Real-world Spatial Question Answering)
ReSQ is a real-world Spatial Question Answering dataset with human-generated questions built on an existing corpus with SpRL annotations.
3 papers · 0 benchmarks
RoFT (Real or Fake Text)
RoFT is a dataset of 21,000 human annotations of generated text.
3 papers · 1 benchmark
RoboPianist is a benchmarking suite for high-dimensional control, targeted at testing high spatial and temporal precision, coordination, and planning, all with an underactuated system frequently making-and-breaking contacts.
3 papers · 0 benchmarks
SAVEE (Surrey Audio-Visual Expressed Emotion)
The Surrey Audio-Visual Expressed Emotion (SAVEE) dataset was recorded as a pre-requisite for the development of an automatic emotion recognition system.
3 papers · 1 benchmark
SCapRepo (Google Play Screenshot Caption)
A screenshot-caption dataset containing 135k pairs of screenshots and captions extracted from Google Play.
3 papers · 0 benchmarks
SCoralDet Dataset (Soft-Coral Detection Dataset)
> High-quality underwater coral detection dataset for machine learning and computer vision research.
3 papers · 1 benchmark
SDN (Situated Dialogue Navigation)
Situated Dialogue Navigation (SDN) is a navigation benchmark of 183 trials with a total of 8415 utterances, around 18.7 hours of control streams, and 2.9 hours of trimmed audio.
3 papers · 0 benchmarks
SIDD-Image (Segmented Intrusion Detection Dataset)
This is the first image-based network intrusion detection dataset.
3 papers · 1 benchmark
SKAB (Skoltech Anomaly Benchmark)
SKAB is designed for evaluating algorithms for anomaly detection.
3 papers · 2 benchmarks
SLNET (SLNET: A Redistributable Corpus of 3rd-party Simulink Models)
SLNET is collection of third party Simulink models.
3 papers · 0 benchmarks
SONICS (Synthetic Or Not - Identifying Counterfeit Songs)
SONICS is a large-scale dataset comprising 97,164 songs — 48,090 real songs from YouTube and 49,074 fake songs from Suno & Udio — designed for synthetic song detection (SSD), also known as fake song detection (FSD).
3 papers · 0 benchmarks
A first-of-its-kind large dataset of sarcastic/non-sarcastic tweets with high-quality labels and extra features: (1) sarcasm perspective labels (2) new contextual features.
3 papers · 0 benchmarks
This dataset is proposed in the ICLR 2024 paper: Measuring Vision-Language STEM Skills of Neural Models.
3 papers · 0 benchmarks
SYNTH-PEDES is a large-scale person dataset with image-text pairs by far, which contains 312,321 identities, 4,791,711 images, and 12,138,157 textual descriptions.
3 papers · 0 benchmarks
SkinCon (SKIN Concepts Dataset)
SkinCon is a skin disease dataset densely annotated by dermatologists.
3 papers · 0 benchmarks
The SmartSpeaker benchmark tests the performance of reacting to music player commands in English as well as in French.
3 papers · 1 benchmark
The dataset is approved for public release, distribution unlimited.
3 papers · 0 benchmarks
Spoken versions of the Semantic Textual Similarity dataset for testing semantic sentence level embeddings.
3 papers · 0 benchmarks
Stanceosaurus is a corpus of 28,033 tweets in English, Hindi, and Arabic annotated with stance towards 251 misinformation claims.
3 papers · 0 benchmarks
This dataset comprises a collection of stellarator configurations used to train the model over multiple iterations.
3 papers · 0 benchmarks
This repository contains a financial-domain-focused dataset for financial sentiment/emotion classification and stock market time series prediction.
3 papers · 0 benchmarks
SupplyGraph (SupplyGraph: A Benchmark Dataset for Supply Chain Planning using Graph Neural Networks)
Graph Neural Networks (GNNs) have gained traction across different domains such as transportation, bio-informatics, language processing, and computer vision.
3 papers · 0 benchmarks
Synthehicle is a massive CARLA-based synthehic multi-vehicle multi-camera tracking dataset and includes ground truth for 2D detection and tracking, 3D detection and tracking, depth estimation, and semantic, instance and panoptic…
3 papers · 1 benchmark
TCAB (Text Classification Attack Benchmark)
Text Classification Attack Benchmark (TCAB) is a dataset for analyzing, understanding, detecting, and labeling adversarial attacks against text classifiers.
3 papers · 0 benchmarks
The TREC News Track features modern search tasks in the news domain.
3 papers · 1 benchmark
The dataset has been designed to represent true web videos in the wild, with good visual quality and diverse content characteristics, The test video collection for TRECVID-AVS2019-TRECVID-AVS2021, which contains 1,082,649 web video clips,…
3 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.