Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 65 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 3073–3120 of 3,998

PWISeg (PWISeg Surgical Instruments Dataset)
Overview The Surgical Instruments Recognition Dataset is a groundbreaking collection of high-resolution images (1280x960 pixels) specifically designed for the recognition and categorization of surgical instruments.
1 paper · 0 benchmarks
PaSa is a dataset to train Machine Learning algorithms to automate the highlighting of patent paragraphs with semantic annotations.
1 paper · 0 benchmarks
1 paper · 0 benchmarks
A dataset of 2D robot recordings with 21 different symbols.
1 paper · 0 benchmarks
We have prepared a dataset, ParagraphOrdreing, which consists of around 300,000 paragraph pairs.
1 paper · 0 benchmarks
The data set includes information about 120+ elections (configuration settings and descriptive statistics), projects and 125k+ anonymized voters and their budget preferences.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
We address the computer-assisted search for prior art by creating a training dataset for supervised machine learning called PatentMatch.
1 paper · 0 benchmarks
PatternCom is a composed image retrieval benchmark based on PatternNet.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Optimization of pedestrian evacuation in different environments
1 paper · 0 benchmarks
Peer to Peer Hate is a comprehensive hate speech dataset capturing various types of hate.
1 paper · 0 benchmarks
Pentachromatic Cultural Palette Dataset is characterized by unique cultural semantics and values.
1 paper · 0 benchmarks
The Perfume Co-Preference Network dataset comprises comprehensive user reviews and ratings collected from the Persian retail platform Atrafshan.
1 paper · 0 benchmarks
The Peripheral Blood Cell} (PBC) dataset consists of 17,092 images.
1 paper · 0 benchmarks
The Permuted bAbi dialog task is an adaptation of the "Dialog bAbI tasks data" dataset released by Facebook.
1 paper · 0 benchmarks
The PEDC is a corpus of 14 episodes of This American Life podcast transcripts that have been annotated for events.
1 paper · 0 benchmarks
PhD (PhD: A ChatGPT-Prompted Visual hallucination Evaluation Dataset)
Multimodal Large Language Models (MLLMs) hallucinate, resulting in an emerging topic of visual hallucination evaluation (VHE).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
PheMT is a phenomenon-wise dataset designed for evaluating the robustness of Japanese-English machine translation systems.
1 paper · 0 benchmarks
Phrase in Context is a curated benchmark for phrase understanding and semantic search, consisting of three tasks of increasing difficulty: Phrase Similarity (PS), Phrase Retrieval (PR) and Phrase Sense Disambiguation (PSD).
1 paper · 0 benchmarks
PhysNLU is a collection of 4 core datasets related to sentence classification, ordering, and coherence of physics explanations based on related tasks.
1 paper · 0 benchmarks
Pirá (Pirá: A Bilingual Portuguese-English Dataset for Question-Answering about the Ocean)
A large set of questions and answers about the ocean and the Brazilian coast both in Portuguese and English.
1 paper · 0 benchmarks
PlainFact is a high-quality human-annotated dataset with fine-grained explanation (i.e., added information) annotations.
1 paper · 0 benchmarks
An evaluation dataset for planning with LLM agents
1 paper · 0 benchmarks
We established a large-scale plant disease segmentation dataset named PlantSeg.
1 paper · 0 benchmarks
Poisoned Water Detection using Smartphone embedded WiFi CSI data and Machine Learning Algorithms (Dataset and machine learning algorithms to detect poisoned water from clean water via using Smartphone embedded Wi-Fi CSI data.)
This repository contains a dataset and machine learning algorithms to detect poisoned water from clean water via using equivalent Smartphone embedded Wi-Fi CSI data.
1 paper · 0 benchmarks
Poker Hand Histories A collection of poker hand histories, covering 11 poker variants, in the poker hand history (PHH) format.
1 paper · 0 benchmarks
The current industrial pipeline includes 315 dynamic industrial scenarios, which can be categorized into three types: QR codes, text, and products.
1 paper · 0 benchmarks
Poly-FEVER is a multilingual fact verification benchmark designed to evaluate hallucination detection in large language models (LLMs).
1 paper · 0 benchmarks
PolyMATH, a challenging benchmark aimed at evaluating the general cognitive reasoning abilities of MLLMs.
1 paper · 0 benchmarks
PolyNews is a multilingual dataset containing news titles in 77 languages and 19 scripts.
1 paper · 0 benchmarks
PolyNews is a multilingual parallel dataset containing news titles 833 language pairs, spanning in 64 languages and 17 scripts.
1 paper · 0 benchmarks
PolyU-BPCoMa (HK PolyU Backpack Colorized Mapping)
PolyU-BPCoMa: A Dataset and Benchmark Towards Mobile Colorized Mapping Using a Backpack Multisensorial System
1 paper · 0 benchmarks
This dataset contains annual Sentinel-2 MSI composites (wet and dry season) for Kigali for the period 2016-2020.
1 paper · 0 benchmarks
Post-hoc Calibration Dataset This dataset collection is designed for the evaluation and development of post-hoc calibration methods for deep neural network classifiers.
1 paper · 0 benchmarks
Pothole Dataset (Pothole_detection_inference.)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
In this Pre-Contest Workshop Slidedeck.pdf: Instructional materials delivered for the seven pre-contest workshops
1 paper · 0 benchmarks
In this Pre-Contest Workshop Video Recordings folder: Seven screen and audio recordings of seven pre-contest workshops
1 paper · 0 benchmarks
PreRAID (Prescreening Rheumatoid Arthritis Information Database (PreRAID))
PreRAID is a structured dataset designed to evaluate the diagnostic capabilities of Large Language Models (LLMs) in Rheumatoid Arthritis (RA) diagnosis.
1 paper · 0 benchmarks
Dataset Description High-level explanation of dataset characteristics: This dataset includes electromyographic (EMG) signals captured using the BiTalino device.
1 paper · 0 benchmarks
ProNCI consists of 22.5K proper noun compounds along with their free-form semantic interpretations.
1 paper · 0 benchmarks
Probability words NLI (Natural language inference with words estimative of probability (WEP))
This dataset tests the capabilities of language models to correctly capture the meaning of words denoting probabilities (WEP), e.g.
1 paper · 1 benchmark
Functionally correct (ok) and incorrect (buggy) solutions to five Probleable Problems: http://arxiv.org/abs/2405.15123 The ok solutions correspond to attempts that successfully probed all ambiguities in the given specification; the buggy…
1 paper · 0 benchmarks
Processed Twitter is a dataset that is used for Twitter topic recognition.
1 paper · 0 benchmarks
The corpus contains review sentences mostly of products in electronics domain, annotated and segregated into 4 comparison categories.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.