Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 33 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 1537–1584 of 3,998

DivEMT (Post-Editing Effort Across Typologically-diverse Languages)
DivEMT, the first publicly available post-editing study of Neural Machine Translation (NMT) over a typologically diverse set of target languages.
3 papers · 0 benchmarks
DrugEHRQA (Electronic Health Record QA)
Contains over 70,000 question-answer pairs from both structured tables and unstructured notes from a publicly available Electronic Health Record (EHR).
3 papers · 0 benchmarks
This repository contains gzipped files containing more than 2 million tokens (words) from answers submitted by more than 6,000 students over the course of their first 30 days of using Duolingo.
3 papers · 0 benchmarks
This is a gzipped CSV file containing the 13 million Duolingo student learning traces used in experiments by Settles & Meeder (2016).
3 papers · 0 benchmarks
E-NER is a publicly available legal Named Entity Recognition (NER) data set.
3 papers · 0 benchmarks
EC-FUNSD is introduced in [[arXiv:2402.02379]](https://arxiv.org/abs/2402.02379) as a benchmark of semantic entity recognition (SER) and entity linking (EL), designed for the entity-centric robustness evaluation of pre-trained…
3 papers · 2 benchmarks
From Grounded Human-Object Interaction Hotspots from Video (ICCV'19): We collect annotations for interaction keypoints on EPIC Kitchens in order to quantitatively evaluate our method in parallel to the OPRA dataset (where annotations are…
3 papers · 1 benchmark
EasyPortrait (Face Parsing and Portrait Segmentation Dataset)
We introduce a large-scale image dataset EasyPortrait for portrait segmentation and face parsing.
3 papers · 0 benchmarks
Ego4D-HCap is a hierarchical video captioning dataset comprised of a three-tier hierarchy of captions: short clip-level captions, medium-length video segment descriptions, and long-range video-level summaries.
3 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 0 benchmarks
EmailSum (Email Thread Summarization)
Email Thread Summarization (EmailSum) is a dataset which contains human-annotated short (<30 words) and long (<100 words) summaries of 2,549 email threads (each containing 3 to 10 emails) over a wide variety of topics.
3 papers · 2 benchmarks
Emotional Dialogue Acts data contains dialogue act labels for existing emotion multi-modal conversational datasets.
3 papers · 0 benchmarks
ErAConD (Error Annotated Conversational Dialog Dataset for Grammatical Error Correction)
ErAConD is a novel GEC dataset consisting of parallel original and corrected utterances drawn from open-domain chatbot conversations.
3 papers · 0 benchmarks
FFHQ-Text is a small-scale face image dataset with large-scale facial attributes, designed for text-to-face generation & manipulation, text-guided facial image manipulation, and other vision-related tasks.
3 papers · 0 benchmarks
FSDD (Free Spoken Digit Dataset)
Free Spoken Digit Dataset (FSDD) is a simple audio/speech dataset consisting of recordings of spoken digits in wav files at 8kHz.
3 papers · 0 benchmarks
We introduce FUNSD-r and CORD-r in Token Path Prediction, the revised VrD-NER datasets to reflect the real-world scenarios of NER on scanned VrDs.
3 papers · 1 benchmark
FakeNewsAMT & Celebrity include two novel datasets for the task of fake news detection, covering seven different news domains.
3 papers · 0 benchmarks
This is a dataset for segmentation and classification of epistemic activities in diagnostic reasoning texts.
3 papers · 0 benchmarks
FireRisk (FireRisk: A Remote Sensing Dataset for Fire Risk Assessment)
In this work, we propose a novel remote sensing dataset, FireRisk, consisting of 7 fire risk classes with a total of 91 872 labelled images for fire risk assessment.
3 papers · 1 benchmark
The Five-Billion-Pixels dataset contains more than 5 billion labeled pixels of 150 high-resolution Gaofen-2 (4 m) satellite images, annotated in a 24-category system covering artificial-constructed, agricultural, and natural classes.
3 papers · 0 benchmarks
FixMyPose is a dataset for automated pose correction.
3 papers · 0 benchmarks
FloDial (Flowchart Grounded Dialogs Dataset)
Flowchart Grounded Dialog Dataset (FloDial) is a corpus of troubleshooting dialogs between a user and an agent collected using Amazon Mechanical Turk.
3 papers · 0 benchmarks
FunQA is a challenging video question answering (QA) dataset specifically designed to evaluate and enhance the depth of video reasoning based on counter-intuitive and fun videos.
3 papers · 0 benchmarks
We present a new large-scale photorealistic panoramic dataset named FutureHouse, which has the following characteristics.
3 papers · 0 benchmarks
GAS (Grasp Area Segmentation)
GAS (Grasp Area Segmentation) dataset consists of 10089 RGB images of cluttered scenes grouped into 1121 grasp-area segmentation tasks.
3 papers · 0 benchmarks
GlassTemp (Glass Transition Temperature)
The GlassTemp dataset is collected from Polyinfo.
3 papers · 1 benchmark
This noisy speech test set is created from the Google Speech Commands v2 [1] and the Musan dataset[2].
3 papers · 1 benchmark
GraSP (Holistic and Multi-Granular Surgical Scene Understanding of Prostatectomies)
Holistic and Multi-Granular Surgical Scene Understanding of Prostatectomies (GraSP) dataset, a curated benchmark that models surgical scene understanding as a hierarchy of complementary tasks with varying levels of granularity.
3 papers · 1 benchmark
GroOT (Grounded Multiple Object Tracking)
One of the recent trends in vision problems is to use natural language captions to describe the objects of interest.
3 papers · 0 benchmarks
HELOC (Home Equity Line of Credit)
HELOC The HELOC dataset from FICO.
3 papers · 1 benchmark
HOPE-Image (Household Objects for Pose Estimation)
The NVIDIA HOPE datasets consist of RGBD images and video sequences with labeled 6-DoF poses for 28 toy grocery objects.
3 papers · 0 benchmarks
HOPE-Video (Household Objects for Pose Estimation)
The HOPE-Video dataset contains 10 video sequences (2038 frames) with 5-20 objects on a tabletop scene captured by a robot arm-mounted RealSense D415 RGBD camera.
3 papers · 0 benchmarks
HSI-Drive is the hyperspectral image (HSI) dataset created by the Digital Electronics Design Group (GDED) of the University of the Basque Country (UPV/EHU).
3 papers · 1 benchmark
HUME-VB (The Hume Vocal Bursts Dataset)
The Hume Vocal Burst Database (H-VB) includes all train, validation, and test recordings and corresponding emotion ratings for the train and validation recordings.
3 papers · 7 benchmarks
HatemojiCheck is a test suite for detecting emoji-based hate of 3,930 test cases covering seven functionalities of emoji-based hate and six identities.
3 papers · 0 benchmarks
Hazards&Robots (Hazards&Robots: A Dataset for Visual Anomaly Detection in Robotics)
We consider the problem of detecting, in the visual sensing data stream of an autonomous mobile robot, semantic patterns that are unusual (i.e., anomalous) with respect to the robot’s previous experience in similar environments.
3 papers · 0 benchmarks
HeadlineCause is a dataset for detecting implicit causal relations between pairs of news headlines.
3 papers · 0 benchmarks
Healthline is a nutrition related dataset for multi-document summarization, using scientific studies.
3 papers · 0 benchmarks
dataset link : https://www.kaggle.com/datasets/osamahosamabdellatif/high-quality-invoice-images-for-ocr Overview High-Quality Invoice Images for OCR is a curated dataset containing professionally scanned and digitally captured invoice…
3 papers · 0 benchmarks
HuRDL (Human-Robot Dialogue Learning Corpus)
The Human-Robot Dialogue Learning (HuRDL) Corpus is a dataset about asking questions in situated task-based interactions.
3 papers · 0 benchmarks
We present datasets containing urban traffic and rural road scenes recorded using hyperspectral snap-shot sensors mounted on a moving car.
3 papers · 1 benchmark
IAM Dataset (A Comprehensive and Large-Scale Dataset for Integrated Argument Mining Tasks)
We introduce a large and comprehensive dataset to facilitate the study of several essential AM tasks in the debating system.
3 papers · 2 benchmarks
ICSI Meeting Corpus in JSON format.
3 papers · 1 benchmark
IIW-400 (ImageInWords: IIW-400)
Please refer: https://github.com/google/imageinwords/blob/main/datasets/IIW-400/README.md
3 papers · 0 benchmarks
IJB-S (IARPA Janus Benchmark-S)
Paper Abstract We present IJB–S dataset, an open-source IARPA Janus Surveillance Video Benchmark and associated protocols.
3 papers · 1 benchmark
We have cleaned the noisy IMDB-WIKI dataset using a constrained clustering method, resulting this new benchmark for in-the-wild age estimation.
3 papers · 1 benchmark
IPAC (Icelandic Parallel Abstracts Corpus)
IPAC (Icelandic Parallel Abstracts Corpus ) is a new Icelandic-English parallel corpus, composed of abstracts from student theses and dissertations.
3 papers · 0 benchmarks
This dataset contains the data for the paper 'Using Multiple Instance Learning for Explainable Solar Flare Prediction'.
3 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.