Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 7 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 289–336 of 3,998
PMC-VQA is a large-scale medical visual question-answering dataset that contains 227k VQA pairs of 149k images that cover various modalities or diseases.
55 papers · 2 benchmarks
QUASAR-T (QUestion Answering by Search And Reading – Trivia)
QUASAR-T is a large-scale dataset aimed at evaluating systems designed to comprehend a natural language query and extract its answer from a large corpus of text.
55 papers · 1 benchmark
Quora Question Pairs (QQP) dataset consists of over 400,000 question pairs, and each question pair is annotated with a binary value indicating whether the two questions are paraphrase of each other.
55 papers · 8 benchmarks
The REVERB (REverberant Voice Enhancement and Recognition Benchmark) challenge is a benchmark for evaluation of automatic speech recognition techniques.
55 papers · 1 benchmark
Jericho is a learning environment for man-made Interactive Fiction (IF) games.
54 papers · 0 benchmarks
The Multi-Domain Sentiment Dataset contains product reviews taken from Amazon.com from many product types (domains).
54 papers · 1 benchmark
Room-Across-Room (RxR) is a multilingual dataset for Vision-and-Language Navigation (VLN) for Matterport3D environments.
54 papers · 1 benchmark
This corpus includes annotations of cancer-related PubMed articles, covering 3 full papers (PMID:24651010, PMID:11777939, PMID:15630473) as well as the result sections of 46 additional PubMed papers.
53 papers · 1 benchmark
DRCD (Delta Reading Comprehension Dataset)
Delta Reading Comprehension Dataset (DRCD) is an open domain traditional Chinese machine reading comprehension (MRC) dataset.
53 papers · 0 benchmarks
Large language models (LLMs), after being aligned with vision models and integrated into vision-language models (VLMs), can bring impressive improvement in image reasoning tasks.
53 papers · 1 benchmark
The ICDAR2003 dataset is a dataset for scene text recognition.
53 papers · 1 benchmark
MLDoc (Multilingual Document Classification Corpus)
Multilingual Document Classification Corpus (MLDoc) is a cross-lingual document classification dataset covering English, German, French, Spanish, Italian, Russian, Japanese and Chinese.
53 papers · 8 benchmarks
The MMVP (Multimodal Visual Patterns) Benchmark focuses on identifying "CLIP-blind pairs" – images that appear similar to the CLIP model despite having clear visual differences.
53 papers · 1 benchmark
PAQ (Probably Asked Questions)
Probably Asked Questions (PAQ) is a very large resource of 65M automatically-generated QA-pairs.
53 papers · 0 benchmarks
EntailmentBank is a dataset that contains multistep entailment trees.
52 papers · 0 benchmarks
InfographicVQA is a dataset that comprises a diverse collection of infographics along with natural language questions and answers annotations.
52 papers · 1 benchmark
The Machine Translation of Noisy Text (MTNT) dataset is a Machine Translation dataset that consists of noisy comments on Reddit and professionally sourced translation.
52 papers · 0 benchmarks
There exist previous works [6, 10] that constructed referring segmentation datasets for videos.
52 papers · 3 benchmarks
WikiLingua includes ~770k article and summary pairs in 18 languages from WikiHow.
52 papers · 1 benchmark
ACE 2004 (ACE 2004 Multilingual Training Corpus)
ACE 2004 Multilingual Training Corpus contains the complete set of English, Arabic and Chinese training data for the 2004 Automatic Content Extraction (ACE) technology evaluation.
51 papers · 6 benchmarks
KuaiRec is a real-world dataset collected from the recommendation logs of the video-sharing mobile app Kuaishou.
51 papers · 0 benchmarks
The large-scale MUSIC-AVQA dataset of musical performance contains 45,867 question-answer pairs, distributed in 9,288 videos for over 150 hours.
51 papers · 1 benchmark
QReCC contains 14K conversations with 81K question-answer pairs.
51 papers · 0 benchmarks
TOFU (Task of Fictitious Unlearning)
The TOFU dataset serves as a benchmark for evaluating unlearning performance of large language models on realistic tasks.
51 papers · 0 benchmarks
Lesion segmentation data includes the original image, paired with the expert manual tracing of the lesion boundaries in the form of a binary mask.
50 papers · 1 benchmark
PlotQA is a VQA dataset with 28.9 million question-answer pairs grounded over 224,377 plots on data from real-world sources and questions based on crowd-sourced question templates.
50 papers · 5 benchmarks
2018 Data Science Bowl (2018 Data Science Bowl Find the nuclei in divergent images to advance medical discovery)
This dataset contains a large number of segmented nuclei images.
49 papers · 1 benchmark
CMRC 2018 (Chinese Machine Reading Comprehension 2018)
CMRC 2018 is a dataset for Chinese Machine Reading Comprehension.
49 papers · 0 benchmarks
DVQA (Data Visualizations via Question Answering)
DVQA is a synthetic question-answering dataset on images of bar-charts.
49 papers · 1 benchmark
Letter (Letter Recognition Data Set)
Letter Recognition Data Set is a handwritten digit dataset.
49 papers · 2 benchmarks
MKQA (Multilingual Knowledge Questions and Answers)
Multilingual Knowledge Questions and Answers (MKQA) is an open-domain question answering evaluation set comprising 10k question-answer pairs aligned across 26 typologically diverse languages (260k question-answer pairs in total).
49 papers · 0 benchmarks
The KIT Motion-Language is a dataset linking human motion and natural language.
48 papers · 2 benchmarks
VOICES (Voices Obscured In Complex Environmental Settings)
The VOICES corpus is a dataset to promote speech and signal processing research of speech recorded by far-field microphones in noisy room conditions.
48 papers · 0 benchmarks
BDD-X (Berkeley Deep Drive-X (eXplanation))
Berkeley Deep Drive-X (eXplanation) is a dataset is composed of over 77 hours of driving within 6,970 videos.
47 papers · 0 benchmarks
ECHR is an English legal judgment prediction dataset of cases from the European Court of Human Rights (ECHR).
47 papers · 1 benchmark
Gait3D is a large-scale 3D representation-based gait recognition dataset.
47 papers · 2 benchmarks
QUASAR (QUestion Answering by Search And Reading)
The Question Answering by Search And Reading (QUASAR) is a large-scale dataset consisting of QUASAR-S and QUASAR-T.
47 papers · 1 benchmark
ScreenSpot Evaluation Benchmark ScreenSpot is an evaluation benchmark for GUI grounding, comprising over 1,200 instructions from various environments, including iOS, Android, macOS, Windows, and Web.
47 papers · 1 benchmark
UAV-Human is a large dataset for human behavior understanding with UAVs.
47 papers · 5 benchmarks
- We present a large and diverse abdominal CT organ segmentation dataset, termed AbdomenCT-1K, with more than 1000 (1K) CT scans from 12 medical centers, including multi-phase, multi-vendor, and multi-disease cases.
46 papers · 0 benchmarks
CoS-E (Commonsense Explanations Dataset)
CoS-E consists of human explanations for commonsense reasoning in the form of natural language sequences and highlighted annotations Source: Explain Yourself!
46 papers · 0 benchmarks
FEVEROUS (Fact Extraction and VERification Over Unstructured and Structured information)
FEVEROUS (Fact Extraction and VERification Over Unstructured and Structured information) is a fact verification dataset which consists of 87,026 verified claims.
46 papers · 0 benchmarks
Legal General Language Understanding Evaluation (LexGLUE) benchmark is a collection of datasets for evaluating model performance across a diverse set of legal NLU tasks in a standardized way.
46 papers · 1 benchmark
OpenLane is the first real-world and the largest scaled 3D lane dataset to date.
46 papers · 2 benchmarks
UDC (Ubuntu Dialogue Corpus)
Ubuntu Dialogue Corpus (UDC) is a dataset containing almost 1 million multi-turn dialogues, with a total of over 7 million utterances and 100 million words.
46 papers · 8 benchmarks
CoSQL (Conversational Text-to-SQL Challenge)
CoSQL is a corpus for building cross-domain, general-purpose database (DB) querying dialogue systems.
45 papers · 1 benchmark
GEM (Generation, Evaluation, and Metrics)
Generation, Evaluation, and Metrics (GEM) is a benchmark environment for Natural Language Generation with a focus on its Evaluation, both through human annotations and automated Metrics.
45 papers · 1 benchmark
A new large-scale geometry problem-solving dataset - 3,002 multi-choice geometry problems - dense annotations in formal language for the diagrams and text - 27,213 annotated diagram logic forms (literals) - 6,293 annotated text logic forms…
45 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.