Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 64 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 3025–3072 of 3,998

Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The dataset identifies the shortcomings of existing benchmarks in evaluating the problem of compositional generalization, which underscores the need for the development of datasets tailored to assess compositional generalization in open…
1 paper · 1 benchmark
Dataset OQRanD and OQGenD for paper "Asking the crowd: Asking the Crowd: Question Analysis, Evaluation and Generation for Open Discussion on Online Forums" by Zi Chai, Xinyu Xing, Xiaojun Wan and Bo Huang.
1 paper · 0 benchmarks
Dataset OQRanD and OQGenD for paper "Asking the crowd: Asking the Crowd: Question Analysis, Evaluation and Generation for Open Discussion on Online Forums" by Zi Chai, Xinyu Xing, Xiaojun Wan and Bo Huang.
1 paper · 0 benchmarks
OSLD (Open Set Logo Detection Dataset)
Open Set Logo Detection Dataset (OSLD Dataset) is a dataset of eCommerce product images with associated brand logo images.
1 paper · 0 benchmarks
This is the paper “DF-RAP: A Robust Adversarial Perturbation for Defending against Deepfakes in Real-world Social Network Scenarios" OSN-transmission CelebA sampling dataset collected by manual upload and download.
1 paper · 0 benchmarks
This dataset contains orthographic samples of words in 19 languages (ar, br, de, en, eno, ent, eo, es, fi, fr, fro, it, ko, nl, pt, ru, sh, tr, zh).
1 paper · 0 benchmarks
OVIC Datasets (Open Vocabulary Image Classification Datasets)
Due to the free-form nature of the open vocabulary image classification task, special annotations are required for image sets used for evaluation purposes.
1 paper · 4 benchmarks
An object-centric version of Stylized COCO to benchmark texture bias and out-of-distribution robustness of vision models.
1 paper · 0 benchmarks
This dataset is a collection of memes from various existing datasets, online forums, and freshly scrapped contents.
1 paper · 0 benchmarks
The Swiss Drone data set was recorded around Cheseaux-sur-Lausanne in Switzerland using a senseFly eBee Classic in 2014 (SenseFly, 2020).
1 paper · 1 benchmark
OllaBench v.0.2 (OllaBench for Interdependent Cybersecurity v.0.2)
Large Language Models (LLMs) have the potential to enhance Agent-Based Modeling by better representing complex interdependent cybersecurity systems, improving cybersecurity threat modeling and risk management.
1 paper · 0 benchmarks
Olympic 2024 is a human-annotated dataset that contains 220 high-quality instance.
1 paper · 0 benchmarks
To effectively evaluate OmniCount across open-vocabulary, supervised, and few-shot counting tasks, a dataset catering to a broad spectrum of visual categories and instances featuring various visual categories with multiple instances and…
1 paper · 2 benchmarks
The OnlySports Dataset is a comprehensive collection of sports-related text data, comprising approximately 600 billion tokens.
1 paper · 0 benchmarks
Dataset of cross-layer Radio Access Network (RAN) Key Performance Measurements (KPMs) and protocol stack logs collected on an Open RAN deployment instantiated on Colosseum with traffic twinned from that of commercial cellular traces.
1 paper · 0 benchmarks
🏃‍♂️ Open-HypermotionX Dataset Open-Hypermotion is a large-scale, high-quality dataset designed for training and evaluating pose-guided human image animation models, with a special focus on complex, dynamic human motions (Hypermotion),…
1 paper · 0 benchmarks
OpenD5 is a a meta-dataset which aggregates 675 open-ended problems ranging across business, social sciences, humanities, machine learning, and health, and uses a set of unified evaluation metrics: validity, relevance, novelty, and…
1 paper · 0 benchmarks
We introduce OpenDebateEvidence, a comprehensive dataset for argument mining and summarization sourced from the American Competitive Debate community.
1 paper · 0 benchmarks
We create the first open-source large-scale S2V generation dataset OpenS2V-5M, which consists of five million high-quality 720P subject-text-video triples.
1 paper · 1 benchmark
Orchid2024 is a fine-grained classification dataset specifically designed for Chinese Cymbidium orchid cultivars.
1 paper · 0 benchmarks
OrdinalDataset (Ordinal Encoding Data set)
It includes 10 data sets that consists of both raw data set and encoded data set where it is encoded through BERT-Sort Encoder with MLM initialization of .
1 paper · 1 benchmark
Overnight is a dataset for semantic parsing in eight domains.
1 paper · 0 benchmarks
The PART-OF dataset is a dataset of relations extracted from a medical ontology.
1 paper · 0 benchmarks
PASSION dataset (PASSION derm 2024 dataset)
Overview PASSION derm is a pioneering initiative dedicated to closing the diversity gap in dermatology datasets.
1 paper · 0 benchmarks
PDFM Embeddings (Population Dynamics Foundation Model Embeddings)
PDFM Embeddings are condensed vector representations designed to encapsulate the complex, multidimensional interactions among human behaviors, environmental factors, and local contexts at specific locations.
1 paper · 0 benchmarks
PECC (PECC: Problem Extraction and Coding Challenges)
Recent advancements in large language models (LLMs) have showcased their exceptional abilities across various tasks, such as code generation, problem-solving and reasoning.
1 paper · 1 benchmark
PEM Fuel Cell Dataset (Proton Exchange Membrane (PEM) Fuel Cell Dataset)
This dataset are about Nafion 112 membrane standard tests and MEA activation tests of PEM fuel cell in various operation condition.
1 paper · 0 benchmarks
PGDataset (Profile Generation Dataset)
PGDataset (Profile Generation Dataset) is a dataset created for the PGTask (Profile Generation Task), where the goal is to extract/generate a profile sentence given a dialogue utterance.
1 paper · 1 benchmark
PHP Webshell Dataset (First PHP Webshell Opcode Incremental Dataset)
First PHP Webshell Opcode Incremental Dataset Motivation To improve the robustness of PHP webshell detection by analyzing low-level opcode patterns, circumventing common code obfuscation and evasion techniques.
1 paper · 0 benchmarks
PIAST (PIAST: A Multimodal Piano Dataset with Audio, Symbolic and Text)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This dataset is well-structured for the physics-informed training of Neural operators for irregular domain geometry, which provides the FEM results of solving a darcy problem in a domain geometry shape of a pentagram.
1 paper · 0 benchmarks
This dataset is well-structured for the physics-informed training of Neural operators for irregular domain geometry, which provides the FEM results of solving a 2D plate stress problem in a domain geometry shape of a rectangle with a hole.
1 paper · 0 benchmarks
We assembled a benchmark of electronic component pinouts, PINS100, containing 100 common parts frequently used in circuits found on high-traffic electronic tutorial websites such as the ARDUINO PROJECT HUB and AUTODESK TINKERCAD CIRCUITS.
1 paper · 0 benchmarks
PIZZA is a dataset for parsing pizza and drink orders, whose semantics cannot be captured by flat slots and intents.
1 paper · 0 benchmarks
PInNED (Personalized Instance-based Navigation Embodied Dataset)
In the last years, the research interest in visual navigation towards objects in indoor environments has grown significantly.
1 paper · 0 benchmarks
PLAD (Point Line and Depth dataset)
PLAD is a dataset where sparse depth is provided by line-based visual SLAM to verify StructMDC.
1 paper · 1 benchmark
PLOD-filtered (PLOD: An Abbreviation Detection Dataset for Scientific Documents)
PLOD: An Abbreviation Detection Dataset This is the PLOD (filtered) Dataset published at LREC 2022.
1 paper · 0 benchmarks
PLOD-unfiltered (PLOD: An Abbreviation Detection Dataset for Scientific Documents)
PLOD: An Abbreviation Detection Dataset This is the PLOD (unfiltered) Dataset published at LREC 2022.
1 paper · 0 benchmarks
PMC-SA (PMC Structured Abstracts)
PMC-SA (PMC Structured Abstracts) is a dataset of academic publications, used for the task of structured summarization.
1 paper · 0 benchmarks
POIE (Products for OCR and Information Extraction)
Products for OCR and Information Extraction (POIE) dataset derives from camera images of various products in the real world.
1 paper · 0 benchmarks
The LiT.RL POLIT-FALSE-n-LEGIT NEWS DB 2016-2017 contains a total of 274 news articles about U.S.
1 paper · 0 benchmarks
The POTUS Corpus is a Database of Weekly Addresses for the Study of Stance in Politics and Virtual Agents.
1 paper · 0 benchmarks
PQ-decaNLP (Paraphrase Questions - decaNLP)
Multitask learning has led to significant advances in Natural Language Processing, including the decaNLP benchmark where question answering is used to frame 10 natural language understanding tasks in a single model.
1 paper · 0 benchmarks
PQAref (Pubmed Question Answering with references)
The PQAref dataset is a dataset for fine-tuning large language models for referenced question-answering in biomedical domain.
1 paper · 0 benchmarks
This is the official dataset for PRMBench.
1 paper · 0 benchmarks
PRONTO (PRONTO heterogeneous benchmark dataset)
The PRONTO heterogeneous benchmark dataset is based on an industrial-scale multiphase flow facility.
1 paper · 1 benchmark
PSB2 (The Second Program Synthesis Benchmark Suite)
https://arxiv.org/abs/2106.06086
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.