Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 80 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 3793–3840 of 3,998

Heel Dataset (Heel Bone X-Ray Dataset)
Heel Bone X-Ray Dataset consists of 3,956 X-ray images of the foot, primarily focused on detecting and classifying heel bone diseases.
0 papers · 0 benchmarks
0 papers · 0 benchmarks
Huawei-UK-University-Challenge-Competition-2021 (Huawei UK University Challenge Competition 2021 - TASK2)
Huawei University Challenge Competition 2021 Data Science for Indoor positioning 2.2 Full Mall Graph Clustering Train The sample training data for this problem is a set of 106981 fingerprints (task2trainfingerprints.json) and some edges…
0 papers · 0 benchmarks
0 papers · 0 benchmarks
Hyper Drive (Hyperspectral Driving Dataset)
Towards automated analysis of large environments, hyperspectral sensors must be adapted into a format where they can be operated from mobile robots.
0 papers · 0 benchmarks
H²O Interaction (Human-to-Human-or-Object Interaction)
H²O is an image dataset annotated for Human-to-human-or-object interaction detection.
0 papers · 0 benchmarks
The International Cardiac Arrest REsearch consortium (I-CARE) Database includes baseline clinical information and continuous electroencephalogram (EEG) and electrocardiogram (ECG) recordings from comatose patients following cardiac arrest.
0 papers · 0 benchmarks
ICConv (A Large-scale Automated Intent-oriented and Context-aware Conversational Search Dataset)
The dataset contains 105,811 information-seeking conversations converted from MS MARCO.
0 papers · 0 benchmarks
Can you detect fraud from customer transactions?
0 papers · 0 benchmarks
Overview The IITKGPFence dataset is designed for tasks related to fence-like occlusion detection, defocus blur, depth mapping, and object segmentation.
0 papers · 0 benchmarks
A large paroxysmal atrial fibrillation long-term electrocardiogram monitoring database Abstract Atrial fibrillation (AF) is the most common sustained heart arrhythmia in adults.
0 papers · 0 benchmarks
This is a Dataset for Arabic/English text detection and optical character recognition.
0 papers · 0 benchmarks
Im-Promptu Visual Analogy Suite is a meta-learning framework.
0 papers · 0 benchmarks
InLUT3D (Indoor Lodz University of Technology Point Cloud Dataset)
This dataset called Indoor Lodz University of Technology Point Cloud Dataset (InLUT3D) is a point cloud set tailored for real object classification and both semantic and instance segmentation tasks.
0 papers · 0 benchmarks
0 papers · 0 benchmarks
This dataset is an extremely challenging set of over 20,000+ original Number plate images captured and crowdsourced from over 700+ urban and rural areas, where each image is manually reviewed and verified by computer vision professionals…
0 papers · 0 benchmarks
0 papers · 0 benchmarks
IndirectRequests is an LLM-generated dataset of user utterances in a task-oriented dialogue setting where the user does not directly specify their preferred slot value.
0 papers · 0 benchmarks
0 papers · 0 benchmarks
InpaintCOCO is a benchmark to understand fine-grained concepts in multimodal models (vision-language) similar to Winoground.
0 papers · 0 benchmarks
0 papers · 0 benchmarks
The Remote Sensing dataset contains the following key features for each annotated marking: Marking Type: Specifies whether the marking is a lane-use arrow or a crosswalk.
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
We introduce the KAIST multi-spectral dataset, which covers a greater range of drivable regions, from urban to residential, for autonomous systems.
0 papers · 0 benchmarks
Dataset Description: Summarized Wiki Articles with TTL Knowledge Graphs Overview This dataset comprises 500 summarized Wikipedia articles, each accompanied by a corresponding TTL knowledge graph.
0 papers · 0 benchmarks
KID-F (K-pop Idol Dataset - Female)
Description K-pop Idol Dataset - Female (KID-F) is the first dataset of K-pop idol high quality face images.
0 papers · 0 benchmarks
KITTI-360-SR (KITTI-360 modification for Scene Recognition task)
Scene Recognition is a problem, where a set of visible objects must be correctly associated with objects marked on a semantic map - this problem is also sometimes called a Data Association.
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
L-SVD (Large-Scale Selfie Video Dataset (L-SVD): A Benchmark for Emotion Recognition)
Welcome to L-SVD L-SVD is an extensive and rigorously curated video dataset aimed at transforming the field of emotion recognition.
0 papers · 0 benchmarks
LMCQA (Legal Multiple Choice Question Answering)
This dataset contains a set of multiple-choice questions related to various legal topics.
0 papers · 0 benchmarks
0 papers · 0 benchmarks
Lemon dataset has been prepared to investigate the possibilities to tackle the issue of fruit quality control.
0 papers · 0 benchmarks
LoLI-Street (Low-Light Images of Streets)
We introduce low-light image enhancement benchmark dataset “Low-light Images of Streets (LoLI-Street),” which contains three subsets: train, validation, and test.
0 papers · 0 benchmarks
Low Light Dataset (Dataset with ill-lighting conditions DILCOD)
Introduced by Khan.
0 papers · 0 benchmarks
The Lusitano dataset was collected over a 3-month period, spanning from January to March, from Paulo de Oliveira, S.A., a prominent textile company, based in Covilhã, Portugal, renowned for its innovative contributions to the textile…
0 papers · 0 benchmarks
0 papers · 0 benchmarks
MAEC (Multimodal Aligned Earnings Conference Call Dataset)
MAEC is a new, large-scale multi-modal, text-audio paired, earnings-call dataset named MAEC, based on S&P 1500 companies.
0 papers · 0 benchmarks
A dataset for multi-context visual grounding.
0 papers · 0 benchmarks
MHRI dataset (Multimodal Human-Robot Interaction dataset)
The dataset includes recordings from 10 different users teaching the robot different common kitchen objects, that consists of synchronized recordings from three cameras and a microphone mounted on the robot: An RGB-d camera covers the user…
0 papers · 0 benchmarks
This dataset contains 4606 articles from 1996 to 2024 that were presented in MIE (Medical Informatics Europe Conference) conferences.
0 papers · 0 benchmarks
MIMIC Meme Dataset (Misogyny Identification in Multimodal Internet Content in Hindi-English Code-Mix Language)
This dataset endeavors to fill the research void by presenting a meticulously curated collection of misogynistic memes in a code-mixed language of Hindi and English.
0 papers · 0 benchmarks
0 papers · 0 benchmarks
IQ testing has served as a foundational methodology for evaluating human cognitive capabilities, deliberately decoupling assessment from linguistic background, language proficiency, or domain-specific knowledge to isolate core competencies…
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.