Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 65 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 3073–3120 of 3,998
PWISeg (PWISeg Surgical Instruments Dataset)
Overview The Surgical Instruments Recognition Dataset is a groundbreaking collection of high-resolution images (1280x960 pixels) specifically designed for the recognition and categorization of surgical instruments.
1 paper · 0 benchmarks
PaSa is a dataset to train Machine Learning algorithms to automate the highlighting of patent paragraphs with semantic annotations.
1 paper · 0 benchmarks
A dataset of 2D robot recordings with 21 different symbols.
1 paper · 0 benchmarks
We have prepared a dataset, ParagraphOrdreing, which consists of around 300,000 paragraph pairs.
1 paper · 0 benchmarks
The data set includes information about 120+ elections (configuration settings and descriptive statistics), projects and 125k+ anonymized voters and their budget preferences.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
We address the computer-assisted search for prior art by creating a training dataset for supervised machine learning called PatentMatch.
1 paper · 0 benchmarks
PatternCom is a composed image retrieval benchmark based on PatternNet.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Optimization of pedestrian evacuation in different environments
1 paper · 0 benchmarks
Peer to Peer Hate is a comprehensive hate speech dataset capturing various types of hate.
1 paper · 0 benchmarks
Pentachromatic Cultural Palette Dataset is characterized by unique cultural semantics and values.
1 paper · 0 benchmarks
The Perfume Co-Preference Network dataset comprises comprehensive user reviews and ratings collected from the Persian retail platform Atrafshan.
1 paper · 0 benchmarks
The Peripheral Blood Cell} (PBC) dataset consists of 17,092 images.
1 paper · 0 benchmarks
The Permuted bAbi dialog task is an adaptation of the "Dialog bAbI tasks data" dataset released by Facebook.
1 paper · 0 benchmarks
The PEDC is a corpus of 14 episodes of This American Life podcast transcripts that have been annotated for events.
1 paper · 0 benchmarks
PhD (PhD: A ChatGPT-Prompted Visual hallucination Evaluation Dataset)
Multimodal Large Language Models (MLLMs) hallucinate, resulting in an emerging topic of visual hallucination evaluation (VHE).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
PheMT is a phenomenon-wise dataset designed for evaluating the robustness of Japanese-English machine translation systems.
1 paper · 0 benchmarks
Phrase in Context is a curated benchmark for phrase understanding and semantic search, consisting of three tasks of increasing difficulty: Phrase Similarity (PS), Phrase Retrieval (PR) and Phrase Sense Disambiguation (PSD).
1 paper · 0 benchmarks
PhysNLU is a collection of 4 core datasets related to sentence classification, ordering, and coherence of physics explanations based on related tasks.
1 paper · 0 benchmarks
Pirá (Pirá: A Bilingual Portuguese-English Dataset for Question-Answering about the Ocean)
A large set of questions and answers about the ocean and the Brazilian coast both in Portuguese and English.
1 paper · 0 benchmarks
PlainFact is a high-quality human-annotated dataset with fine-grained explanation (i.e., added information) annotations.
1 paper · 0 benchmarks
An evaluation dataset for planning with LLM agents
1 paper · 0 benchmarks
We established a large-scale plant disease segmentation dataset named PlantSeg.
1 paper · 0 benchmarks
This repository contains a dataset and machine learning algorithms to detect poisoned water from clean water via using equivalent Smartphone embedded Wi-Fi CSI data.
1 paper · 0 benchmarks
Poker Hand Histories A collection of poker hand histories, covering 11 poker variants, in the poker hand history (PHH) format.
1 paper · 0 benchmarks
The current industrial pipeline includes 315 dynamic industrial scenarios, which can be categorized into three types: QR codes, text, and products.
1 paper · 0 benchmarks
Poly-FEVER is a multilingual fact verification benchmark designed to evaluate hallucination detection in large language models (LLMs).
1 paper · 0 benchmarks
PolyMATH, a challenging benchmark aimed at evaluating the general cognitive reasoning abilities of MLLMs.
1 paper · 0 benchmarks
PolyNews is a multilingual dataset containing news titles in 77 languages and 19 scripts.
1 paper · 0 benchmarks
PolyNews is a multilingual parallel dataset containing news titles 833 language pairs, spanning in 64 languages and 17 scripts.
1 paper · 0 benchmarks
PolyU-BPCoMa: A Dataset and Benchmark Towards Mobile Colorized Mapping Using a Backpack Multisensorial System
1 paper · 0 benchmarks
This dataset contains annual Sentinel-2 MSI composites (wet and dry season) for Kigali for the period 2016-2020.
1 paper · 0 benchmarks
Post-hoc Calibration Dataset This dataset collection is designed for the evaluation and development of post-hoc calibration methods for deep neural network classifiers.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
In this Pre-Contest Workshop Slidedeck.pdf: Instructional materials delivered for the seven pre-contest workshops
1 paper · 0 benchmarks
In this Pre-Contest Workshop Video Recordings folder: Seven screen and audio recordings of seven pre-contest workshops
1 paper · 0 benchmarks
PreRAID (Prescreening Rheumatoid Arthritis Information Database (PreRAID))
PreRAID is a structured dataset designed to evaluate the diagnostic capabilities of Large Language Models (LLMs) in Rheumatoid Arthritis (RA) diagnosis.
1 paper · 0 benchmarks
Dataset Description High-level explanation of dataset characteristics: This dataset includes electromyographic (EMG) signals captured using the BiTalino device.
1 paper · 0 benchmarks
ProNCI consists of 22.5K proper noun compounds along with their free-form semantic interpretations.
1 paper · 0 benchmarks
This dataset tests the capabilities of language models to correctly capture the meaning of words denoting probabilities (WEP), e.g.
1 paper · 1 benchmark
Functionally correct (ok) and incorrect (buggy) solutions to five Probleable Problems: http://arxiv.org/abs/2405.15123 The ok solutions correspond to attempts that successfully probed all ambiguities in the given specification; the buggy…
1 paper · 0 benchmarks
Processed Twitter is a dataset that is used for Twitter topic recognition.
1 paper · 0 benchmarks
The corpus contains review sentences mostly of products in electronics domain, annotated and segregated into 4 comparison categories.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.