Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 11 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 481–528 of 3,998
DUDE (Document UnderstanDing of Everything)
DUDE is formulated as an instance of Document Question Answering (DocQA) to evaluate how well current solutions deal with multi-page documents, if they can navigate and reason over the layout, and if they can generalize these skills to…
26 papers · 0 benchmarks
Head and Neck Tumor Segmentation
26 papers · 0 benchmarks
Dataset is constructed from single intent dataset SNIPS.
26 papers · 2 benchmarks
MultiDoc2Dial (MultiDoc2Dial: Modeling Dialogues Grounded in Multiple Documents)
MultiDoc2Dial is a new task and dataset on modeling goal-oriented dialogues grounded in multiple documents.
26 papers · 0 benchmarks
🤖 Robo3D - The SemanticKITTI-C Benchmark SemanticKITTI-C is an evaluation benchmark heading toward robust and reliable 3D semantic segmentation in autonomous driving.
26 papers · 1 benchmark
VALSE (VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena)
We propose VALSE (Vision And Language Structured Evaluation), a novel benchmark designed for testing general-purpose pretrained vision and language (V&L) models for their visio-linguistic grounding capabilities on specific linguistic…
26 papers · 12 benchmarks
WikiReading is a large-scale natural language understanding task and publicly-available dataset with 18 million instances.
26 papers · 0 benchmarks
XStoryCloze consists of the professionally translated version of the English StoryCloze dataset (Spring 2016 version) to 10 non-English languages.
26 papers · 0 benchmarks
BDD-A (Berkeley DeepDrive Attention)
Dataset Statistics: The statistics of our dataset are summarized and compared with the largest existing dataset (DR(eye)VE) [1] in Table 1.
25 papers · 0 benchmarks
BioRED is a first-of-its-kind biomedical relation extraction dataset with multiple entity types (e.g.
25 papers · 3 benchmarks
DIOR-RSVG is a large-scale benchmark dataset of remote sensing data (RSVG).
25 papers · 0 benchmarks
ELEVATER (Evaluation of Language-augmented Visual Task-level Transfer)
The ELEVATER benchmark is a collection of resources for training, evaluating, and analyzing language-image models on image classification and object detection.
25 papers · 2 benchmarks
HONEST (Hurtful Sentence Completion in English Language Models)
The HONEST dataset is a template-based corpus for testing the hurtfulness of sentence completions in language models (e.g., BERT) in six different languages (English, Italian, French, Portuguese, Romanian, and Spanish).
25 papers · 1 benchmark
LargeST (LargeST: A Benchmark Dataset for Large-Scale Traffic Forecasting)
In this work, we propose LargeST as a new benchmark dataset (see Figure 1), with the goal of facilitating the development of accurate and efficient methods in the context of large-scale traffic forecasting.
25 papers · 1 benchmark
The MULTEXT-East resources are a multilingual dataset for language engineering research and development.
25 papers · 0 benchmarks
OASST1 (OpenAssistant Conversations Dataset)
license: apache-2.0 tags: human-feedback sizecategories: 100K Languages with under 1000 messages Vietnamese: 952 Basque: 947 Polish: 886 Hungarian: 811 Arabic: 666 Dutch: 628 Swedish: 512 Turkish: 454 Finnish: 386 Czech: 372 Danish: 358…
25 papers · 0 benchmarks
TOPv2 (Task Oriented Parsing v2)
Task Oriented Parsing v2 (TOPv2) representations for intent-slot based dialog systems.
25 papers · 0 benchmarks
A benchmark dataset for the Aspect Sentiment Triplet Extraction, an updated version of ASTE-Data-V1.
24 papers · 1 benchmark
FlickrStyle10K is collected and built on Flickr30K image caption dataset.
24 papers · 2 benchmarks
GeoS is a dataset for automatic math problem solving.
24 papers · 1 benchmark
InterHuman is a multimodal dataset, named InterHuman.
24 papers · 1 benchmark
The Segmenting and Tracking Every Pixel (STEP) benchmark consists of 21 training sequences and 29 test sequences.
24 papers · 2 benchmarks
MMSE-HR (Multimodal Spontaneous Expression-Heart Rate dataset)
The MMSE-HR benchmark consists of a dataset of 102 videos from 40 subjects recorded at 1040x1392 raw resolution at 25fps.
24 papers · 1 benchmark
Moral Stories is a crowd-sourced dataset of structured narratives that describe normative and norm-divergent actions taken by individuals to accomplish certain intentions in concrete situations, and their respective consequences.
24 papers · 0 benchmarks
The Natural Stories dataset consists of English texts edited to contain many low-frequency syntactic constructions while still sounding fluent to native speakers.
24 papers · 0 benchmarks
REALY (Region-aware benchmark based on the LYHM)
The REALY benchmark aims to introduce a region-aware evaluation pipeline to measure the fine-grained normalized mean square error (NMSE) of 3D face reconstruction methods from under-controlled image sets.
24 papers · 2 benchmarks
SIMMC (Situated and Interactive Multimodal Conversations)
Situated Interactive MultiModal Conversations (SIMMC) is the task of taking multimodal actions grounded in a co-evolving multimodal input content in addition to the dialog history.
24 papers · 0 benchmarks
News translation is a recurring WMT task.
24 papers · 0 benchmarks
The Wiki-ZSL (Wiki Zero-Shot Learning) dataset contains 113 relations and 94,383 instances from Wikipedia.
24 papers · 1 benchmark
node classification on twitch-gamers
24 papers · 2 benchmarks
4D-DRESS (A 4D Dataset of Real-world Human Clothing with Semantic Annotations)
4D-DRESS is the first real-world 4D dataset of human clothing, capturing 64 human outfits in more than 520 motion sequences.
23 papers · 4 benchmarks
ARCTIC (Articulated Objects in Free-form Hand Interaction)
ARCTIC is a dataset of free-form interactions of hands and articulated objects.
23 papers · 0 benchmarks
We propose EMAGE, a framework to generate full-body human gestures from audio and masked gestures, encompassing facial, local body, hands, and global movements.
23 papers · 2 benchmarks
DocUNet (Document Image Unwarping via a Stacked U-Net)
Various documents dataset.
23 papers · 3 benchmarks
HPS (Human POSEitioning System Dataset)
HPS Dataset is a collection of 3D humans interacting with large 3D scenes (300-1000 m², up to 2500 m²).
23 papers · 0 benchmarks
HRSOD (High-Resolution Salient Object Detection)
There exist several datasets for saliency detection, but none of them is specifically designed for high-resolution salient object detection.
23 papers · 1 benchmark
OpenImages V6 is a large-scale dataset , consists of 9 million training images, 41,620 validation samples, and 125,456 test samples.
23 papers · 2 benchmarks
PIE-Bench (Prompt-based Image Editing Benchmark)
PIE-Bench comprises 700 images featuring 10 distinct editing types.
23 papers · 1 benchmark
PointCloud-C is the very first test-suite for point cloud robustness analysis under corruptions.
23 papers · 2 benchmarks
RECCON is a dataset for the task of recognizing emotion cause in conversations.
23 papers · 2 benchmarks
RST-DT (RST Discourse Treebank)
The Rhetorical Structure Theory (RST) Discourse Treebank consists of 385 Wall Street Journal articles from the Penn Treebank annotated with discourse structure in the RST framework along with human-generated extracts and abstracts…
23 papers · 2 benchmarks
Real 3D-AD is the first point cloud anomaly detection dataset for industrial products.
23 papers · 2 benchmarks
ShapeWorld is a new evaluation methodology and framework for multimodal deep learning models, with a focus on formal-semantic style generalization capabilities.
23 papers · 0 benchmarks
Toyota Smarthome Trimmed has been designed for the activity classification task of 31 activities.
23 papers · 0 benchmarks
514 algebra word problems and associated equation systems gathered from Algebra.com.
22 papers · 1 benchmark
AMR Bank (Abstract Meaning Representation)
The AMR Bank is a set of English sentences paired with simple, readable semantic representations.
22 papers · 1 benchmark
The Easy Communications (EasyCom) dataset is a world-first dataset designed to help mitigate the cocktail party effect from an augmented-reality (AR) -motivated multi-sensor egocentric world view.
22 papers · 4 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.