Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 9 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 385–432 of 3,998
Doc2Dial (Doc2Dial: Document-grounded Dialogue)
For goal-oriented document-grounded dialogs, it often involves complex contexts for identifying the most relevant information, which requires better understanding of the inter-relations between conversations and documents.
36 papers · 0 benchmarks
EBM-NLP annotates PICO (Participants, Interventions, Comparisons and Outcomes) spans in clinical trial abstracts.
36 papers · 1 benchmark
A hand-object interaction dataset with 3D pose annotations of hand and object.
36 papers · 2 benchmarks
In this project, we introduce InfoSeek, a visual question answering dataset tailored for information-seeking questions that cannot be answered with only common sense knowledge.
36 papers · 2 benchmarks
LSUI (Large Scale Underwater Image Dataset)
We released a large-scale underwater image (LSUI) dataset including 5004 image pairs, which involve richer underwater scenes (lighting conditions, water types and target categories) and better visual quality reference images than the…
36 papers · 1 benchmark
MAD (Movie Audio Descriptions) is an automatically curated large-scale dataset for the task of natural language grounding in videos or natural language moment retrieval.
36 papers · 2 benchmarks
PanLex translates words in thousands of languages.
36 papers · 0 benchmarks
TEACh (Task-driven Embodied Agents that Chat)
Robots operating in human spaces must be able to engage in natural language interaction with people, both understanding and executing instructions, and using conversation to resolve ambiguity and recover from mistakes.
36 papers · 0 benchmarks
A large-scale V2X perception dataset using CARLA and OpenCDA
36 papers · 1 benchmark
VisualMRC (VisualMRC: Machine Reading Comprehension on Document Images)
VisualMRC is a visual machine reading comprehension dataset that proposes a task: given a question and a document image, a model produces an abstractive answer.
36 papers · 1 benchmark
Attribution, Relation, and Order (ARO) benchmark to systematically evaluate the ability of VLMs to understand different types of relationships, attributes, and order information.
35 papers · 0 benchmarks
CIRCO (Composed Image Retrieval on Common Objects in context)
CIRCO (Composed Image Retrieval on Common Objects in context) is an open-domain benchmarking dataset for Composed Image Retrieval (CIR) based on real-world images from COCO 2017 unlabeled set.
35 papers · 1 benchmark
This dataset contains complex tables from the annual reports of S&P 500 companies with detailed table structure annotations to help table structure recognition and table data extraction.
35 papers · 0 benchmarks
The Open Entity dataset is a collection of about 6,000 sentences with fine-grained entity types annotations.
35 papers · 2 benchmarks
CoVoST is a large-scale multilingual speech-to-text translation corpus.
34 papers · 0 benchmarks
The DIHARD II development and evaluation sets draw from a diverse set of sources exhibiting wide variation in recording equipment, recording environment, ambient noise, number of speakers, and speaker demographics.
34 papers · 1 benchmark
GrailQA (Strongly Generalizable Question Answering)
GrailQA is a new large-scale, high-quality dataset for question answering on knowledge bases (KBQA) on Freebase with 64,331 questions annotated with both answers and corresponding logical forms in different syntax (i.e., SPARQL,…
34 papers · 4 benchmarks
Re-DocRED (Revisiting Document Level Relation Extraction)
The Re-DocRED Dataset resolved the following problems of DocRED: 1.
34 papers · 3 benchmarks
2000 HUB5 English Evaluation Transcripts was developed by the Linguistic Data Consortium (LDC) and consists of transcripts of 40 English telephone conversations used in the 2000 HUB5 evaluation sponsored by NIST (National Institute of…
33 papers · 2 benchmarks
The Dialog State Tracking Challenges 2 & 3 (DSTC2&3) were research challenge focused on improving the state of the art in tracking the state of spoken dialog systems.
33 papers · 5 benchmarks
GLUCOSE is a large-scale dataset of implicit commonsense causal knowledge, encoded as causal mini-theories about the world, each grounded in a narrative context.
33 papers · 0 benchmarks
We build a large-scale, comprehensive, and high-quality synthetic dataset for city-scale neural rendering researches.
33 papers · 0 benchmarks
We describe the SemEval task of extracting keyphrases and relations between them from scientific documents, which is crucial for understanding which publications describe which processes, tasks and materials.
33 papers · 1 benchmark
WMCA (Wide Multi Channel Presentation Attack)
The Wide Multi Channel Presentation Attack (WMCA) database consists of 1941 short video recordings of both bonafide and presentation attacks from 72 different identities.
33 papers · 1 benchmark
WMT 2020 is a collection of datasets used in shared tasks of the Fifth Conference on Machine Translation.
33 papers · 0 benchmarks
WikiEvents is a document-level event extraction benchmark dataset which includes complete event and coreference annotation.
33 papers · 1 benchmark
🤖 Robo3D - The nuScenes-C Benchmark nuScenes-C is an evaluation benchmark heading toward robust and reliable 3D perception in autonomous driving.
33 papers · 2 benchmarks
CholecT50 is a dataset of endoscopic videos of laparoscopic cholecystectomy surgery introduced to enable research on fine-grained action recognition in laparoscopic surgery.
32 papers · 5 benchmarks
Dress Code is a new dataset for image-based virtual try-on composed of image pairs coming from different catalogs of YOOX NET-A-PORTER.
32 papers · 1 benchmark
Electricity (Individual household electric power consumption Data Set)
Abstract: Measurements of electric power consumption in one household with a one-minute sampling rate over a period of almost 4 years.
32 papers · 6 benchmarks
EmoBank is a corpus of 10k English sentences balancing multiple genres, annotated with dimensional emotion metadata in the Valence-Arousal-Dominance (VAD) representation format.
32 papers · 0 benchmarks
IPM NEL (Derczynski IPM Named Entity Linking)
This data is for the task of named entity recognition and linking/disambiguation over tweets.
32 papers · 1 benchmark
MS-CXR (Making the Most of Text Semantics to Improve Biomedical Vision-Language Processing)
The MS-CXR dataset provides 1162 image–sentence pairs of bounding boxes and corresponding phrases, collected across eight different cardiopulmonary radiological findings, with an approximately equal number of pairs for each finding.
32 papers · 0 benchmarks
MedQuAD (Medical Question Answering Dataset)
MedQuAD includes 47,457 medical question-answer pairs created from 12 NIH websites (e.g.
32 papers · 0 benchmarks
ModelNet40-C is a comprehensive dataset to benchmark the corruption robustness of 3D point cloud recognition.
32 papers · 2 benchmarks
P-Stance: A Large Dataset for Stance Detection in Political Domain 2021
32 papers · 1 benchmark
PACO (Parts and Attributes of Common Objects)
Parts and Attributes of Common Objects (PACO) is a detection dataset that goes beyond traditional object boxes and masks and provides richer annotations such as part masks and attributes.
32 papers · 0 benchmarks
PUBHEALTH is a comprehensive dataset for explainable automated fact-checking of public health claims.
32 papers · 0 benchmarks
Worldtree is a corpus of explanation graphs, explanatory role ratings, and associated tablestore.
32 papers · 0 benchmarks
The Extreme Summarization (XSum) dataset is a dataset for evaluation of abstractive single-document summarization systems.
32 papers · 5 benchmarks
The Actor-Action Dataset (A2D) by Xu et al.
31 papers · 1 benchmark
ArtEmis is a large-scale dataset aimed at providing a detailed understanding of the interplay between visual content, its emotional effect, and explanations for the latter in language.
31 papers · 0 benchmarks
CAVE (Multispectral imaging using multiplexed illumination.)
Multispectral imaging using multiplexed illumination.
31 papers · 1 benchmark
ContractNLI is a dataset for document-level natural language inference (NLI) on contracts whose goal is to automate/support a time-consuming procedure of contract review.
31 papers · 0 benchmarks
IGLUE (Image-Grounded Language Understanding Evaluation)
The Image-Grounded Language Understanding Evaluation (IGLUE) benchmark brings together—by both aggregating pre-existing datasets and creating new ones—visual question answering, cross-modal retrieval, grounded reasoning, and grounded…
31 papers · 0 benchmarks
The Image Paragraph Captioning dataset allows researchers to benchmark their progress in generating paragraphs that tell a story about an image.
31 papers · 1 benchmark
The IMAGE-CHAT dataset is a large collection of (image, style trait for speaker A, style trait for speaker B, dialogue between A & B) tuples that we collected using crowd-workers, Each dialogue consists of consecutive turns by speaker A…
31 papers · 2 benchmarks
KVQA (Knowledge-aware VQA)
It contains manually verified 183K question-answer pairs about more than 18K persons and 24K images.
31 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.