Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 13 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 577–624 of 3,998
FoodSeg103 is a new food image dataset containing 7,118 images.
19 papers · 1 benchmark
Funcom is a collection of ~2.1 million Java methods and their associated Javadoc comments.
19 papers · 0 benchmarks
InfiMM-Eval (Complex Open-ended Reasoning Evaluation for Multi-Modal Language Models)
Multi-modal Large Language Models (MLLMs) are increasingly prominent in the field of artificial intelligence.
19 papers · 1 benchmark
Consists of annotated frames containing GI procedure tools such as snares, balloons and biopsy forceps, etc.
19 papers · 3 benchmarks
MCubeS (Multimodal Material Segmentation Dataset)
Multimodal material segmentation (MCubeS) dataset contains 500 sets of images from 42 street scenes.
19 papers · 1 benchmark
NNE is a dataset for Nested Named Entity Recognition in English Newswire
19 papers · 1 benchmark
OVEN (Open-domain Visual Entity Recognition)
In this project, we formally present the task of Open-domain Visual Entity recognitioN (OVEN), where a model need to link an image onto a Wikipedia entity with respect to a text query.
19 papers · 1 benchmark
PKLot (A Robust Dataset for Parking Lot Classification)
The PKLot dataset contains 12,417 images of parking lots and 695,899 images of parking spaces segmented from them, which were manually checked and labeled.
19 papers · 1 benchmark
QAMPARI is an ODQA benchmark, where question answers are lists of entities, spread across many paragraphs.
19 papers · 0 benchmarks
A dataset on asking Questions for Lack of Clarity in open-domain information-seeking conversations.
19 papers · 0 benchmarks
SUTD-TrafficQA (Singapore University of Technology and Design - Traffic Question Answering) is a dataset which takes the form of video QA based on 10,080 in-the-wild videos and annotated 62,535 QA pairs, for benchmarking the cognitive…
19 papers · 1 benchmark
There are now many computer programs for automatically determining the sense of a word in context (Word Sense Disambiguation or WSD).
19 papers · 0 benchmarks
Node classification on Squirrel with the fixed 48%/32%/20% splits provided by Geom-GCN.
19 papers · 2 benchmarks
Node classification on Squirrel with 60%/20%/20% random splits for training/validation/test.
19 papers · 1 benchmark
Taskmaster-1 is a dialog dataset consisting of 13,215 task-based dialogs in English, including 5,507 spoken and 7,708 written dialogs created with two distinct procedures.
19 papers · 0 benchmarks
Torque is an English reading comprehension benchmark built on 3.2k news snippets with 21k human-generated questions querying temporal relationships.
19 papers · 1 benchmark
UMLS (Unified Medical Language System)
The Unified Medical Language System (UMLS) is a comprehensive resource that integrates and disseminates essential terminology, classification standards, and coding systems.
19 papers · 1 benchmark
Contains data from three platforms, i.e., synthetic drones, satellites and ground cameras of 1,652 university buildings around the world.
19 papers · 2 benchmarks
The rounD dataset introduces a fresh compilation of natural road user trajectory data from German roundabouts, gathered using drone technology to navigate past usual challenges such as occlusions inherent in traditional traffic data…
19 papers · 0 benchmarks
This dataset includes reviews (ratings, text, helpfulness votes), product metadata (descriptions, category information, price, brand, and image features), and links (also viewed/also bought graphs).
18 papers · 3 benchmarks
We collect a new dataset of human-posed free-form natural language questions about CLEVR images.
18 papers · 1 benchmark
DWIE (Deutsche Welle corpus for Information Extraction)
The 'Deutsche Welle corpus for Information Extraction' (DWIE) is a multi-task dataset that combines four main Information Extraction (IE) annotation sub-tasks: (i) Named Entity Recognition (NER), (ii) Coreference Resolution, (iii) Relation…
18 papers · 5 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
18 papers · 1 benchmark
The Endomapper dataset is the first collection of complete endoscopy sequences acquired during regular medical practice, including slow and careful screening explorations, making secondary use of medical data.
18 papers · 0 benchmarks
This dataset was collected and prepared by the CALO Project (A Cognitive Assistant that Learns and Organizes).
18 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
18 papers · 1 benchmark
Groningen Meaning Bank is a semantic resource that anyone can edit and that integrates various semantic phenomena, including predicate-argument structure, scope, tense, thematic roles, animacy, pronouns, and rhetorical relations.
18 papers · 0 benchmarks
Subjective video quality assessment (VQA) strongly depends on semantics, context, and the types of visual distortions.
18 papers · 1 benchmark
LIVECell (Label-free In Vitro image Examples of Cells)
The LIVECell (Label-free In Vitro image Examples of Cells) dataset is a large-scale microscopic image dataset for instance-segmentation of individual cells in 2D cell cultures.
18 papers · 1 benchmark
LSHTC is a dataset for large-scale text classification.
18 papers · 0 benchmarks
MED (Monotonicity Entailment Dataset)
MED is a new evaluation dataset that covers a wide range of monotonicity reasoning that was created by crowdsourcing and collected from linguistics publications.
18 papers · 1 benchmark
Nam (A holistic approach to cross-channel image noise modeling and its application to image denoising)
A holistic approach to cross-channel image noise modeling and its application to image denoising
18 papers · 1 benchmark
Node classification on PubMed with 60%/20%/20% random splits for training/validation/test.
18 papers · 1 benchmark
SALAD-Bench (A Hierarchical and Comprehensive Safety Benchmark for Large Language Models)
In the rapidly evolving landscape of Large Language Models (LLMs), ensuring robust safety measures is paramount.
18 papers · 0 benchmarks
SCAND (Socially CompliAnt Navigation Dataset)
Have you wondered how autonomous mobile robots should share space with humans in public spaces?
18 papers · 0 benchmarks
SCICAP is a large-scale image captioning dataset that contains real-world scientific figures and captions.
18 papers · 1 benchmark
SeaDronesSee (SeaDronesSee: A Maritime Benchmark for Detecting Humans in Open Water)
SeaDronesSee is a large-scale data set aimed at helping develop systems for Search and Rescue (SAR) using Unmanned Aerial Vehicles (UAVs) in maritime scenarios.
18 papers · 3 benchmarks
TuringBench is a benchmark environment that contains : - Benchmark tasks- Turing Test (i.e., human vs.
18 papers · 2 benchmarks
TempEval-3 (TempEval-3: events, times, and temporal relations)
Within the SemEval-2013 evaluation exercise, the TempEval-3 shared task aims to advance research on temporal information processing.
18 papers · 2 benchmarks
The first parallel corpus composed from United Nations documents published by the original data creator.
18 papers · 0 benchmarks
VidSitu is a dataset for the task of semantic role labeling in videos (VidSRL).
18 papers · 0 benchmarks
Violin (VIdeO-and-Language INference)
Video-and-Language Inference is the task of joint multimodal understanding of video and text.
18 papers · 0 benchmarks
Many existing datasets for lidar place recognition are solely representative of structured urban environments, and have recently been saturated in performance by deep learning based approaches.
18 papers · 1 benchmark
Node classification on Wisconsin with 60%/20%/20% random splits for training/validation/test.
18 papers · 1 benchmark
X-CSQA is a multilingual dataset for Commonsense reasoning research, based on CSQA.
18 papers · 0 benchmarks
ZInd (Zillow Indoor Dataset)
The Zillow Indoor Dataset (ZInD) provides extensive visual data that covers a real world distribution of unfurnished residential homes.
18 papers · 1 benchmark
xSID (Cross-lingual Slot and Intent Detection)
xSID, a new evaluation benchmark for cross-lingual (X) Slot and Intent Detection in 13 languages from 6 language families, including a very low-resource dialect, covering Arabic (ar), Chinese (zh), Danish (da), Dutch (nl), English (en),…
18 papers · 0 benchmarks
AVeriTeC (AVeriTeC: A Dataset for Real-world Claim Verification with Evidence from the Web)
AVeriTeC (Automated Verification of Textual Claims) is a dataset of 4568 real-world claims covering fact-checks by 50 different organizations.
17 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.