Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 14 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 625–672 of 3,998
BUG is a large-scale gender bias dataset of 108K diverse real-world English sentences, sampled semiautomatically from large corpora using lexical syntactic pattern matching
17 papers · 0 benchmarks
CLEVR-Ref+ is a synthetic diagnostic dataset for referring expression comprehension.
17 papers · 1 benchmark
Node classification on Chameleon with 60%/20%/20% random splits for training/validation/test.
17 papers · 2 benchmarks
ConvQuestions is the first realistic benchmark for conversational question answering over knowledge graphs.
17 papers · 0 benchmarks
DialoGLUE is a natural language understanding benchmark for task-oriented dialogue designed to encourage dialogue research in representation-based transfer, domain adaptation, and sample-efficient task learning.
17 papers · 2 benchmarks
EgoTask QA benchmark contains 40K balanced question-answer pairs selected from 368K programmatically generated questions generated over 2K egocentric videos.
17 papers · 1 benchmark
FreebaseQA is a data set for open-domain QA over the Freebase knowledge graph.
17 papers · 0 benchmarks
By perturbing the widely used GSM8K dataset, an adversarial dataset for grade-school math called GSM-Plus is created.
17 papers · 1 benchmark
Kleister NDA is a dataset for Key Information Extraction (KIE).
17 papers · 1 benchmark
MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities
17 papers · 1 benchmark
MMDialog is a large-scale multi-turn dialogue dataset containing multi-modal open-domain conversations derived from real human-human chat content in social media.
17 papers · 1 benchmark
MoCA-Mask (Moving Camouflaged Animals (MoCA)-Mask)
The original Moving Camouflaged Animals (MoCA) Dataset includes 37K frames from 141 YouTube Video sequences with resolution and sampling rate of 720 × 1280 and 24fps in the majority of cases.
17 papers · 1 benchmark
OpenXAI is the first general-purpose lightweight library that provides a comprehensive list of functions to systematically evaluate the quality of explanations generated by attribute-based explanation methods.
17 papers · 0 benchmarks
A large-scale English paraphrase dataset that surpasses prior work in both quantity and quality.
17 papers · 0 benchmarks
PartialSpoof is a dataset of partially-spoofed data to evaluate detection of partially-spoofed speech data.
17 papers · 0 benchmarks
QAMR (Question-Answer Meaning Representation Dataset)
Question-Answer Meaning Representation (QAMR) represents a predicate-argument structure of a sentence with a set of question-answer pairs, so that annotations can be easily provided by non-experts.
17 papers · 0 benchmarks
Recent accelerations in multi-modal applications have been made possible with the plethora of image and text data available online.
17 papers · 0 benchmarks
Quasimodo is commonsense knowledge base that focuses on salient properties of objects.
17 papers · 0 benchmarks
SECOND (SEmantic Change detectiON Dataset)
SECOND is a well-annotated semantic change detection dataset.
17 papers · 1 benchmark
SODA is a high-quality social dialogue dataset.
17 papers · 0 benchmarks
How to capture the present knowledge from surrounding situations and perform reasoning accordingly is crucial and challenging for machine intelligence.
17 papers · 2 benchmarks
The SUN-SEG dataset is a high-quality per-frame annotated VPS dataset, which includes 158,690 frames from the famous SUN dataset.
17 papers · 1 benchmark
The SUN-SEG dataset is a high-quality per-frame annotated VPS dataset, which includes 158,690 frames from the famous SUN dataset.
17 papers · 1 benchmark
Automated source code generation is currently a popular machine learning-based task.
17 papers · 0 benchmarks
The Switchboard-1 Telephone Speech Corpus (LDC97S62) consists of approximately 260 hours of speech and was originally collected by Texas Instruments in 1990-1, under DARPA sponsorship.
17 papers · 1 benchmark
VQA-E is a dataset for Visual Question Answering with Explanation, where the models are required to generate and explanation with the predicted answer.
17 papers · 0 benchmarks
A temporal counterfactual dataset composing of 1000 short and natural video-caption pairs.
17 papers · 1 benchmark
A publicly available dataset with 242k labeled sections in English and German from two distinct domains: diseases and cities.
17 papers · 0 benchmarks
iSarcasm is a dataset of tweets, each labelled as either sarcastic or nonsarcastic.
17 papers · 1 benchmark
CMU-MOSI (Multimodal Corpus of Sentiment Intensity)
The Multimodal Corpus of Sentiment Intensity (CMU-MOSI) dataset is a collection of 2199 opinion video clips.
16 papers · 2 benchmarks
Casual Conversations dataset is designed to help researchers evaluate their computer vision and audio models for accuracy across a diverse set of age, genders, apparent skin tones and ambient lighting conditions.
16 papers · 0 benchmarks
CholecT45 is a subset of CholecT50 consisting of 45 videos from the Cholec80 dataset.
16 papers · 1 benchmark
Node classification on Cornell with the fixed 48%/32%/20% splits provided by Geom-GCN.
16 papers · 2 benchmarks
Node classification on Cornell with 60%/20%/20% random splits for training/validation/test.
16 papers · 2 benchmarks
DiPCo (DiPCo -- Dinner Party Corpus)
We present a speech data corpus that simulates a "dinner party" scenario taking place in an everyday home environment.
16 papers · 0 benchmarks
FaithDial is a new benchmark for hallucination-free dialogues, by editing hallucinated responses in the Wizard of Wikipedia (WoW) benchmark.
16 papers · 0 benchmarks
GRIT (General Robust Image Task Benchmark)
The General Robust Image Task (GRIT) Benchmark is an evaluation-only benchmark for evaluating the performance and robustness of vision systems across multiple image prediction tasks, concepts, and data sources.
16 papers · 5 benchmarks
GitTables is a corpus of currently 1M relational tables extracted from CSV files in GitHub covering 96 topics.
16 papers · 0 benchmarks
ICFG-PEDES (Identity-Centric and Fine-Grained Person Description Dataset)
One large-scale database for Text-to-Image Person Re-identification, i.e., Text-based Person Retrieval.
16 papers · 3 benchmarks
IDRiD (Indian Diabetic Retinopathy Image Dataset)
Indian Diabetic Retinopathy Image Dataset (IDRiD) dataset consists of typical diabetic retinopathy lesions and normal retinal structures annotated at a pixel level.
16 papers · 3 benchmarks
MATRES (Multi-Axis Temporal RElations for Start-points)
This is the Multi-Axis Temporal RElations for Start-points (i.e., MATRES) dataset
16 papers · 2 benchmarks
MathBench (MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark)
MathBench is an All in One math dataset for language model evaluation, with: A Sophisticated Five-Stage Difficulty Mechanism: Unlike the usual mathematical datasets that can only evaluate a single difficulty level or have a mix of unclear…
16 papers · 0 benchmarks
MeetingBank, a benchmark dataset created from the city councils of 6 major U.S.
16 papers · 1 benchmark
NewsCLIPpings is a dataset for detecting mismatched images and captions.
16 papers · 0 benchmarks
OntoNotes Release 4.0 contains the content of earlier releases -- OntoNotes Release 1.0 LDC2007T21, OntoNotes Release 2.0 LDC2008T04 and OntoNotes Release 3.0 LDC2009T24 -- and adds newswire, broadcast news, broadcast conversation and web…
16 papers · 1 benchmark
SPGISpeech (pronounced “speegie-speech”) is a large-scale transcription dataset, freely available for academic research.
16 papers · 1 benchmark
SPoC (Pseudocode-to-Code)
Pseudocode-to-Code (SPoC) is a program synthesis dataset, containing 18,356 programs with human-authored pseudocode and test cases.
16 papers · 2 benchmarks
TV show Caption is a large-scale multimodal captioning dataset, containing 261,490 caption descriptions paired with 108,965 short video moments.
16 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.