Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 34 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 1585–1632 of 3,998

IfAct (Identifying Human Actions Visible in Online Vlogs)
We consider the task of identifying human actions visible in online videos.
3 papers · 0 benchmarks
IllusionVQA is a Visual Question Answering (VQA) dataset with two sub-tasks.
3 papers · 2 benchmarks
Imgur5k is a large-scale handwritten in-the-wild dataset, containing challenging real world handwritten samples from nearly 5K writers.
3 papers · 0 benchmarks
Information Extraction from Tables (Extraction materials compositions from tables of materials science research papers)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 0 benchmarks
Itihasa is a large-scale corpus for Sanskrit to English translation containing 93,000 pairs of Sanskrit shlokas and their English translations.
3 papers · 1 benchmark
JUSTICE (JUSTICE: A Dataset for Supreme Court’s Judgment Prediction)
The dataset contains 3304 cases from the Supreme Court of the United States from 1955 to 2021.
3 papers · 0 benchmarks
The Jamendo Corpus is a voice detection dataset consisting of 93 songs with Creative Commons license from the Jamendo free music sharing website.
3 papers · 0 benchmarks
Kvasir-VQA (A Text-Image Pair GI Tract Dataset)
The Kvasir-VQA dataset is an extended dataset derived from the HyperKvasir and Kvasir-Instrument datasets, augmented with question-and-answer annotations.
3 papers · 0 benchmarks
The dataset contains a Video capsule endoscopy dataset for polyp segmentation.
3 papers · 1 benchmark
LAGENDA (Layer Age and Gender Dataset)
The LAGENDA dataset is a large-scale dataset with age and gender annotations for face and body bounding boxes.
3 papers · 4 benchmarks
LAION-COCO is the world’s largest dataset of 600M generated high-quality captions for publicly available web-images.
3 papers · 1 benchmark
LARC (Language-annotated Abstraction and Reasoning)
LARC is a dataset built from ARC (Abstraction and Reasoning Corpus).
3 papers · 0 benchmarks
A Large Dataset for Remote Sensing Image Change Captioning.
3 papers · 0 benchmarks
The dataset was proposed in LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.
3 papers · 0 benchmarks
Got "pubchemsmilescanonical.zip" from https://ibm.ent.box.com/v/MoLFormer-data
3 papers · 0 benchmarks
LibriVoxDeEn is a corpus of sentence-aligned triples of German audio, German text, and English translation, based on German audiobooks.
3 papers · 0 benchmarks
MCVQA (Multilingual and Code-mixed Visual Question Answering)
The MCVQA dataset consists of 248, 349 training questions and 121, 512 validation questions for real images in Hindi and Code-mixed.
3 papers · 0 benchmarks
MDID (Multimodal Document Intent Dataset)
The Multimodal Document Intent Dataset (MDID) is a dataset for computing author intent from multimodal data from Instagram.
3 papers · 0 benchmarks
The MG-ShopDial dataset contains English conversations that mix different conversational goals, including search, recommendation, and question answering in the domain of e-commerce.
3 papers · 0 benchmarks
MIMIC-IV ICD-10 contains 122,279 discharge summaries—free-text medical documents—annotated with ICD-10 diagnosis and procedure codes.
3 papers · 1 benchmark
Question Answering (QA) is a widely-used framework for developing and evaluating an intelligent machine.
3 papers · 0 benchmarks
MISP2021 (Multimodal Information Based Speech Processing 2021)
The MISP2021 challenge dataset is a collection of audio-visual conversational data recorded in a home TV scenario using distant multi-microphones.
3 papers · 0 benchmarks
Operating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assistance, situational awareness, and patient safety.
3 papers · 3 benchmarks
MMCode is a multi-modal code generation dataset designed to evaluate the problem-solving skills of code language models in visually rich contexts (i.e.
3 papers · 0 benchmarks
The main goal of the data collection is to acquire highly natural conversations that cover a wide variety of styles and scenarios.
3 papers · 2 benchmarks
MOMA-LRG (Multi-Object Multi-Actor activity parsing with Language-Refined Graphs)
A dataset dedicated to multi-object, multi-actor activity parsing.
3 papers · 1 benchmark
MRS (Multilingual Reply Suggestion)
MRS, a multilingual reply suggestion dataset with ten languages.
3 papers · 0 benchmarks
MS-FIMU (Mobility Scenario FIMU)
Open Dataset: Mobility Scenario FIMU An open, multidimensional (6 categorical attributes), and synthetic dataset of faked virtual humans generated by an optimization approach applied to a real-life call-detail-records-based anonymized…
3 papers · 0 benchmarks
Frame-to-frame video alignment/synchronization
3 papers · 1 benchmark
MUC-4 (Fourth Message Uunderstanding Conference)
A dataset for evaluate system's understanding of given passages.
3 papers · 1 benchmark
The “Medico automatic polyp segmentation challenge” aims to develop computer-aided diagnosis systems for automatic polyp segmentation to detect all types of polyps (for example, irregular polyp, smaller or flat polyps) with high efficiency…
3 papers · 1 benchmark
MeltingTemp (Melting Temperature)
The MeltingTemp dataset is collected from Polyinfo.
3 papers · 0 benchmarks
We sample 2025 frames of images from the original KITTI for Mono3DRefer, containing 41,140 expressions in total and a vocabulary of 5,271 words.
3 papers · 0 benchmarks
The dataset, comprising 1204 meticulously curated images, serves as a comprehensive resource for advancing real-time mosquito detection models.
3 papers · 0 benchmarks
MuLD (Multitask Long Document Benchmark)
MuLD (Multitask Long Document Benchmark) is a set of 6 NLP tasks where the inputs consist of at least 10,000 words.
3 papers · 5 benchmarks
Multilingual automatic speech recognition (ASR) in the medical domain serves as a foundational task for various downstream applications such as speech translation, spoken language understanding, and voice-activated assistants.
3 papers · 0 benchmarks
Multilingual TOP is a dataset for multilingual semantic parsing with human-written sentences as opposed to machine translated ones.
3 papers · 0 benchmarks
NarraSum is a large-scale narrative summarization dataset.
3 papers · 0 benchmarks
Natural Hazards is a natural disaster dataset with sentiment labels, which contains nearly 50,00 Twitter data about different natural disasters in the United States (e.g., a tornado in 2011, a hurricane named Sandy in 2012, a series of…
3 papers · 0 benchmarks
The dataset contains Amazon products from 10 product categories with full human annotations.
3 papers · 2 benchmarks
OREBA (Objectively Recognizing Eating Behavior and Associated Intake)
The OREBA dataset aims to provide a comprehensive multi-sensor recording of communal intake occasions for researchers interested in automatic detection of intake gestures.
3 papers · 0 benchmarks
Evaluate radar localization in diverse environments Download: https://drive.google.com/drive/folders/1uATfrAe-KHlz29e-Ul8qUbUKwPxBFIhP Download
3 papers · 0 benchmarks
OpenCHAIR is a benchmark for evaluating open-vocabulary hallucinations in image captioning models.
3 papers · 0 benchmarks
OpenLane-V2 is the world's first perception and reasoning benchmark for scene structure in autonomous driving.
3 papers · 1 benchmark
OSAI introduces OpenTTGames - an open dataset aimed at evaluation of different computer vision tasks in Table Tennis: ball detection, semantic segmentation of humans, table and scoreboard and fast in-game events spotting.
3 papers · 0 benchmarks
OpenTrench3D, the first publicly available point cloud dataset of underground utilities from open trenches.
3 papers · 1 benchmark
A benchmark designed to evaluate MLLMs’ proficiency in understanding inter-object relationships and textual content.
3 papers · 0 benchmarks
P3 (Psychophysical Patterns Dataset)
A set of patterns used in psychophysical research to evaluate the ability of saliency algorithms to find targets distinct from distractors in orientation, color and size.
3 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.