Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 23 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 1057–1104 of 3,998

RefMatte (Referring Image Matting)
RefMatte is the first large-scale challenging dataset under the task referring image matting, generated by a comprehensive image composition and expression generation engine on top of current public high-quality matting foregrounds with…
7 papers · 3 benchmarks
A data set containing citations, citation contexts, and papers.
7 papers · 0 benchmarks
Understanding spatial relations (e.g., “laptop on table”) in visual input is important for both humans and robots.
7 papers · 1 benchmark
RuBQ (Russian Knowledge Base Questions)
The first Russian knowledge base question answering (KBQA) dataset.
7 papers · 0 benchmarks
English subset of the SLAKE dataset, comprising 642 images and more than 7,000 question–answer pairs.
7 papers · 0 benchmarks
SME (Standard Multimodal Explanation)
SME is a new dataset for Multi-modal Explanation for Visual Question Answering comprising 1,028,230 samples, with 1,656 visual objects requiring detection in explanations.
7 papers · 1 benchmark
SNARE, short for ShapeNet Annotated with Referring Expressions, is a benchmark requires a model to choose which of two objects is being referenced by a natural language description.
7 papers · 0 benchmarks
SODA-A is a large-scale benchmark specialized for small object detection task under aerial scenes, which has 800203 instances with oriented rectangle box annotation across 9 classes.
7 papers · 0 benchmarks
SSC (Spiking Speech Commands v0.2)
The SSC dataset is a spiking version of the Speech Commands dataset release by Google (Speech Commands).
7 papers · 1 benchmark
A multimodal agent benchmark on professional data science and engineering.
7 papers · 0 benchmarks
The Sunnybrook Cardiac Data (SCD), also known as the 2009 Cardiac MR Left Ventricle Segmentation Challenge data, consist of 45 cine-MRI images from a mixed of patients and pathologies: healthy, hypertrophy, heart failure with infarction…
7 papers · 0 benchmarks
The TACRED-Revisited dataset improves the crowd-sourced TACRED dataset for relation extraction by relabeling the dev and test sets using expert linguistic annotators.
7 papers · 1 benchmark
THEODORE (Learning from THEODORE)
Recent work about synthetic indoor datasets from perspective views has shown significant improvements of object detection results with Convolutional Neural Networks(CNNs).
7 papers · 0 benchmarks
TIAGE is a topic-shift aware dialog benchmark constructed utilizing human annotations on topic shifts.
7 papers · 0 benchmarks
TransCG is the first large-scale real-world dataset for transparent object depth completion and grasping, which contains 57,715 RGB-D images of 51 transparent objects and many opaque objects captured from different perspectives (~240…
7 papers · 1 benchmark
UPAR (Unified Pedestrian Attribute Recognition)
The Task: The challenge will use an extension of the UPAR Dataset [1], which consists of images of pedestrians annotated for 40 binary attributes.
7 papers · 1 benchmark
This dataset was collected with the goal of assessing dialog evaluation metrics.
7 papers · 1 benchmark
This dataset was collected with the goal of assessing dialog evaluation metrics.
7 papers · 1 benchmark
VNHSGE (VietNamese High School Graduation Examination Dataset for Large Language Models)
The VNHSGE (VietNamese High School Graduation Examination) dataset, developed exclusively for evaluating large language models (LLMs), is introduced in this article.
7 papers · 9 benchmarks
VQA-CE (VQA Counterexamples)
This dataset provides a new split of VQA v2 (similarly to VQA-CP v2), which is built of questions that are hard to answer for biased models.
7 papers · 1 benchmark
VidHOI is a video-based human-object interaction detection benchmark.
7 papers · 2 benchmarks
A multivariate spatio-temporal benchmark dataset for meteorological forecasting based on real-time observation data from ground weather stations.
7 papers · 16 benchmarks
WikiDetox (Wikipedia Detox)
An annotated dataset of 1m crowd-sourced annotations that cover 100k talk page diffs (with 10 judgements per diff) for personal attacks, aggression, and toxicity.
7 papers · 0 benchmarks
k-qa (K-QA: A Real-World Medical Q&A Benchmark)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
7 papers · 0 benchmarks
A Modular Simulation Framework and Benchmark for Robot Learning.
7 papers · 0 benchmarks
AH36M (Ambiguous Human3.6M)
Since H36M is captured in a controlled environment, it rarely depicts challenging real-world scenarios such as body occlusions that are the main source of ambiguity in the single-view 3D shape estimation problem.
6 papers · 1 benchmark
The AMI Meeting Corpus is a multi-modal data set comprising 100 hours of meeting recordings.
6 papers · 1 benchmark
ARCADE (Automatic Region-based Coronary Artery Disease diagnostics using x-ray angiography imagEs Dataset)
ARCADE: Automatic Region-based Coronary Artery Disease diagnostics using x-ray angiography imagEs Dataset Phase 2 consist of two folders with 300 images in each of them as well as annotations.
6 papers · 0 benchmarks
ASCEND (A Spontaneous Chinese-English Dataset) introduces a high-quality resource of spontaneous multi-turn conversational dialogue Chinese code-switching corpus collected in Hong Kong.
6 papers · 0 benchmarks
Adult Census Income (adult_census_income)
This data was extracted from the 1994 Census bureau database by Ronny Kohavi and Barry Becker (Data Mining and Visualization, Silicon Graphics).
6 papers · 1 benchmark
This dataset was created using a dataset used for data categorization that onsists of 2225 documents from the BBC news website corresponding to stories in five topical areas from 2004-2005 used in the paper of D.
6 papers · 0 benchmarks
BC4CHEMD (BioCreative IV Chemical compound and drug name recognition)
Introduced by Krallinger et al.
6 papers · 1 benchmark
BIPIA (Benchmark of Indirect Prompt Injection Attacks)
Recent advancements in large language models (LLMs) have led to their adoption across various applications, notably in combining LLMs with external content to generate responses.
6 papers · 0 benchmarks
BMELD is a bilingual (English-Chinese) dialogue corpus for Neural chat translation.
6 papers · 0 benchmarks
BiPaR is a manually annotated bilingual parallel novel-style machine reading comprehension (MRC) dataset, developed to support monolingual, multilingual and cross-lingual reading comprehension on novels.
6 papers · 0 benchmarks
CMD is a publicly available collection of hundreds of thousands 2D maps and 3D grids containing different properties of the gas, dark matter, and stars from more than 2,000 different universes.
6 papers · 0 benchmarks
CLAMS (Cross-linguistic Analysis of Models on Syntax)
Targeted syntactic evaluation datasets in 5 languages: English, French, German, Russian, and Hebrew.
6 papers · 0 benchmarks
A dataset of 12-lead ECGs with annotations.
6 papers · 1 benchmark
Median house prices for California districts derived from the 1990 census.
6 papers · 2 benchmarks
CoDa (The Color Dataset)
The Color Dataset (CoDa) is a probing dataset to evaluate the representation of visual properties in language models.
6 papers · 0 benchmarks
Concepticon (Concepticon. A Resource for the Linking of Concept Lists)
This resource, our Concepticon, links concept labels from different conceptlists to concept sets.
6 papers · 0 benchmarks
ConvoSumm is a suite of four datasets to evaluate a model’s performance on a broad spectrum of conversation data.
6 papers · 0 benchmarks
DSC (10 tasks) (Task Incremental Document Sentiment Classification)
A set of 10 DSC datasets (reviews of 10 products) to produce sequences of tasks.
6 papers · 1 benchmark
DebateSum consists of 187328 debate documents, arguments (also can be thought of as abstractive summaries, or queries), word-level extractive summaries, citations, and associated metadata organized by topic-year.
6 papers · 1 benchmark
DocCVQA (Document Collection Visual Question Answering)
DocCVQA is a Document Visual Question Answering dataset, where the questions are posed over a whole collection of 14,362 scanned documents.
6 papers · 0 benchmarks
DroneSURF (DroneSURF: Benchmark Dataset for Drone-based Face Recognition)
Drone Surveillance of Faces, is a large-scale drone dataset intended to facilitate research for face recognition using drones.
6 papers · 1 benchmark
EarthVQA (A multi-modal multi-task VQA dataset for remote sensing)
Earth vision research typically focuses on extracting geospatial object locations and categories but neglects the exploration of relations between objects and comprehensive reasoning.
6 papers · 1 benchmark
EmoDB Dataset (Berlin Database of Emotional Speech)
The EMODB database is the freely available German emotional database.
6 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.