Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 31 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 1441–1488 of 3,998
Relative Human (RH) contains multi-person in-the-wild RGB images with rich human annotations, including: Depth layers: relative depth relationship/ordering between all people in the image.
4 papers · 2 benchmarks
The Restaurant-ACOS dataset is constructed based on the SemEval 2016 Restaurant dataset (Pontiki et al., 2016) and its expansion datasets (Fan et al., 2019; Xu et al., 2020).
4 papers · 1 benchmark
RetVQA (Retrieval-Based Visual Question Answering)
The RetVQA dataset is a large-scale dataset designed for Retrieval-Based Visual Question Answering (RetVQA).
4 papers · 1 benchmark
RidgeBase (RidgeBase: A Cross-Sensor Multi-Finger Contactless Fingerprint Dataset)
Contactless fingerprint matching using smartphone cameras can alleviate major challenges of traditional fingerprint systems including hygienic acquisition, portability and presentation attacks.
4 papers · 0 benchmarks
Text and signs around roads provide crucial information for drivers, vital for safe navigation and situational awareness.
4 papers · 1 benchmark
SCDE is a human-created sentence cloze dataset, collected from public school English examinations in China.
4 papers · 1 benchmark
The Situated Corpus Of Understanding Transactions (SCOUT) is a multi-modal collection of human-robot dialogue in the task domain of collaborative exploration.
4 papers · 0 benchmarks
SDSD-outdoor (Seeing Dynamic Scene in the Dark: High-Quality Video Dataset with Mechatronic Alignment)
Seeing Dynamic Scene in the Dark: High-Quality Video Dataset with Mechatronic Alignment
4 papers · 1 benchmark
SPACE is a large-scale opinion summarization benchmark for the evaluation of unsupervised summarizers.
4 papers · 1 benchmark
SV-Ident (Survey Variable Identification)
SV-Ident comprises 4,248 sentences from social science publications in English and German.
4 papers · 2 benchmarks
SYMON (Synopses of Movie Narratives)
Contains 5,193 video summaries of popular movies and TV series.
4 papers · 0 benchmarks
SciDuet is a dataset for training and benchmarking models for automating document-to-slides generation.
4 papers · 0 benchmarks
SecQA is a specialized dataset created for the evaluation of Large Language Models (LLMs) in the domain of computer security.
4 papers · 0 benchmarks
Separated COCO is automatically generated subsets of COCO val dataset, collecting separated objects for a large variety of categories in real images in a scalable manner, where target object segmentation mask is separated into distinct…
4 papers · 1 benchmark
SoftAttributes (SoftAttributes: Relative movie attribute dataset for soft attributes)
The dataset consists of sets of movie titles, with each set annotated with a single English soft attribute (subjective descriptive property, such as 'confusing' or 'romantic') and a reference movie.
4 papers · 0 benchmarks
Solar-Power (Solar Power Data for Integration Studies (Alabama))
Solar Power Data for Integration Studies NREL's Solar Power Data for Integration Studies are synthetic solar photovoltaic (PV) power plant data points for the United States representing the year 2006.
4 papers · 2 benchmarks
Something-Something-100 is a dataset split created from Something-Something V2.
4 papers · 1 benchmark
SynthPAI (SynthPAI: A Synthetic Dataset for Personal Attribute Inference)
SynthPAI was created to provide a dataset that can be used to investigate the personal attribute inference (PAI) capabilities of LLM on online texts.
4 papers · 1 benchmark
Our goal is to enable deep learning research in neuroscience by releasing the largest publicly available unencumbered database of EEG recordings.
4 papers · 1 benchmark
TemporalWiki is a lifelong benchmark for ever-evolving LMs that utilizes the difference between consecutive snapshots of English Wikipedia and English Wikidata for training and evaluation, respectively.
4 papers · 0 benchmarks
TexBiG (from the German Text-Bild-Gefüge, meaning Text-Image-Structure) is a document layout analysis dataset for historical documents in the late 19th and early 20th century.
4 papers · 4 benchmarks
> The data we use include 366 monthly series, 427 quarterly series and 518 yearly series.
4 papers · 0 benchmarks
This is an entity-level Twitter Sentiment Analysis dataset.
4 papers · 1 benchmark
USIS10K (Large-scale Underwater Salient Instance Segmentation Dataset)
We construct the first large-scale dataset, USIS10K, for the underwater salient instance segmentation task, which contains 10,632 images and pixel-level annotations of 7 categories.
4 papers · 0 benchmarks
UniProtQA consists of proteins and textual queries about their functions and properties.
4 papers · 1 benchmark
VANiLLa is a dataset for Question Answering over Knowledge Graphs (KGQA) offering answers in natural language sentences.
4 papers · 0 benchmarks
Visuelle 2.0 is a dataset containing real data for 5355 clothing products of the retail fast-fashion Italian company, Nuna Lie.
4 papers · 2 benchmarks
A Large Vision-Language Model Knowledge Editing Benchmark
4 papers · 0 benchmarks
The VideoNavQA dataset contains pairs of questions and videos generated in the House3D environment.
4 papers · 0 benchmarks
VirtualHome2KG is a system for constructing and augmenting knowledge graphs (KGs) of daily living activities using virtual space.
4 papers · 0 benchmarks
The VizWiz-VQA-Grounding dataset is a dataset that visually grounds answers to visual questions asked by people with visual impairments.
4 papers · 0 benchmarks
WEC-eng is a cross-document event coreference resolution dataset extracted from English Wikipedia.
4 papers · 0 benchmarks
WHU-RS19 is a set of satellite images exported from Google Earth, which provides high-resolution satellite images up to 0.5 m.
4 papers · 0 benchmarks
WikiCLIR is a large-scale (German-English) retrieval data set for Cross-Language Information Retrieval (CLIR).
4 papers · 0 benchmarks
WikiGraphs is a dataset of Wikipedia articles each paired with a knowledge graph, to facilitate the research in conditional text generation, graph generation and graph representation learning.
4 papers · 1 benchmark
The WikiSem500 dataset contains around 500 per-language cluster groups for English, Spanish, German, Chinese, and Japanese (a total of 13,314 test cases).
4 papers · 0 benchmarks
SRL is the task of extracting semantic predicate-argument structures from sentences.
4 papers · 0 benchmarks
XED is a multilingual fine-grained emotion dataset.
4 papers · 0 benchmarks
XWINO is a multilingual collection of Winograd Schemas in six languages that can be used for evaluation of cross-lingual commonsense reasoning capabilities.
4 papers · 1 benchmark
eICU-CRD (eICU Collaborative Research Database)
The eICU Collaborative Research Database is a large multi-center critical care database made available by Philips Healthcare in partnership with the MIT Laboratory for Computational Physiology.
4 papers · 2 benchmarks
This dataset contains 1304 de-identified longitudinal medical records describing 296 patients.
4 papers · 1 benchmark
t4d (Thinking is for Doing)
This dataset was generated by the code implementation found here: https://github.com/sachith-gunasekara/t4d
4 papers · 0 benchmarks
The dataset is designed specifically to solve a range of computer vision problems (2D-3D tracking, posture) faced by biologists while designing behavior studies with animals.
3 papers · 0 benchmarks
Robot grasping is often formulated as a learning problem.
3 papers · 0 benchmarks
This dataset gathers 10,874 title and abstract pairs from the ACL Anthology Network (until 2016).
3 papers · 1 benchmark
Attention Deficit Hyperactivity Disorder (ADHD) affects at least 5-10% of school-age children and is associated with substantial lifelong impairment, with annual direct costs exceeding $36 billion/year in the US.
3 papers · 0 benchmarks
The AND Dataset contains 13700 handwritten samples and 15 corresponding expert examined features for each sample.
3 papers · 1 benchmark
ARCH2S (Dataset, Benchmark for Learning Exterior Architectural Structures from Point Clouds)
Precise segmentation of architectural structures provides detailed information about various building components, enhancing our understanding and interaction with our built environment.
3 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.