Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 19 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 865–912 of 3,998
MACS (Multi-Annotator Captioned Soundscapes)
This is a dataset containing audio captions and corresponding audio tags for a number of 3930 audio files of the TAU Urban Acoustic Scenes 2019 development dataset (airport, public square, and park).
10 papers · 0 benchmarks
MP-DocVQA (Multipage Document Visual Question Answering)
The dataset is aimed to perform Visual Question Answering on multipage industry scanned documents.
10 papers · 0 benchmarks
MUGEN is a large-scale video-audio-text dataset MUGEN, collected using the open-sourced platform game CoinRun.
10 papers · 0 benchmarks
MUStARD++ is a multimodal sarcasm detection dataset (MUStARD) pre-annotated with 9 emotions.
10 papers · 1 benchmark
MeshRIR is a dataset of acoustic room impulse responses (RIRs) at finely meshed grid points.
10 papers · 0 benchmarks
A large-scale dataset that consists of 21,184 claims, where each claim is assigned a truthfulness label and ruling statement, with 58,523 pieces of evidence in the form of text and images.
10 papers · 0 benchmarks
- NeoRL is a collection of environments and datasets for offline reinforcement learning with a special focus on real-world applications.
10 papers · 0 benchmarks
OCW (Only Connect Wall Dataset and creative problem solving tasks)
The OCW dataset is for evaluating creative problem solving tasks by curating the problems and human performance results from the popular British quiz show Only Connect.
10 papers · 1 benchmark
OLIVES Dataset (Ophthalmic Labels for Investigating Visual Eye Semantics)
Clinical diagnosis of the eye is performed over multifarious data modalities including scalar clinical labels, vectorized biomarkers, two-dimensional fundus images, and three-dimensional Optical Coherence Tomography (OCT) scans.
10 papers · 0 benchmarks
PLAsTiCC (Photometric LSST Astronomical Time-Series Classification Challenge)
The PLAsTiCC dataset is a collection of simulated light curves from the Photometric LSST Astronomical Time-Series Classification Challenge.
10 papers · 0 benchmarks
PointQA is a set of datasets for Visual Question Datasets (VQA) that require a pointer to an object in the image to be answered correctly.
10 papers · 0 benchmarks
REFLACX (Reports and eye-tracking data for localization of abnormalities in chest x-rays)
The REFLACX dataset contains eye-tracking data for 3,032 readings of chest x-rays by five radiologists.
10 papers · 0 benchmarks
REFinD (REFinD: Relation Extraction Financial Dataset)
REFinD is a large-scale annotated dataset of relations, with ∼29K instances and 22 relations amongst 8 types of entity pairs, generated entirely over financial documents.
10 papers · 0 benchmarks
Rainbow is multi-task benchmark for common-sense reasoning that uses different existing QA datasets: aNLI, Cosmos QA, HellaSWAG.
10 papers · 0 benchmarks
ReQA (Retrieval Question-Answering)
Retrieval Question-Answering (ReQA) benchmark tests a model’s ability to retrieve relevant answers efficiently from a large set of documents.
10 papers · 0 benchmarks
ReaSCAN (ReaSCAN: Compositional Reasoning in Language Grounding)
ReaSCAN is a synthetic navigation task that requires models to reason about surroundings over syntactically difficult languages.
10 papers · 0 benchmarks
SST-3 (Stanford Sentiment Treebank: 3-way)
SST-5 is the Stanford Sentiment Treebank 5-way classification dataset (positive, somewhat positive, neutral, somewhat negative, negative).
10 papers · 1 benchmark
SkillSpan (Hard and Soft Skill Extraction from English Job Postings)
SkillSpan is a dataset for Skill Extraction (SE).
10 papers · 0 benchmarks
SpaceNet 2: Building Detection v2 - is a dataset for building footprint detection in geographically diverse settings from very high resolution satellite images.
10 papers · 1 benchmark
First large-scale symphony generation dataset.
10 papers · 1 benchmark
TMED (Tufts Medical Echocardiogram Dataset)
TMED is a clinically-motivated benchmark dataset for computer vision and machine learning from limited labeled data.
10 papers · 0 benchmarks
ViP-Bench (Making Large Multimodal Models Understand Arbitrary Visual Prompts)
ViP-Bench is a comprehensive benchmark designed to assess the capability of multimodal models in understanding visual prompts across multiple dimensions.
10 papers · 1 benchmark
e-ViL is a benchmark for explainable vision-language tasks.
10 papers · 0 benchmarks
The Sixth Informatics for Integrating Biology and the Bedside (i2b2) Natural Language Processing Challenge for Clinical Records focused on the temporal relations in clinical narratives.
9 papers · 2 benchmarks
Our dataset which consists of multiple indoor and outdoor experiments for up to 30 m gNB-UE link.
9 papers · 0 benchmarks
ART consists of over 20k commonsense narrative contexts and 200k explanations.
9 papers · 0 benchmarks
The detection and localization of highly realistic deepfake audio-visual content are challenging even for the most advanced state-of-the-art methods.
9 papers · 0 benchmarks
Is an acronym disambiguation (AD) dataset for scientific domain with 62,441 samples which is significantly larger than the previous scientific AD dataset.
9 papers · 0 benchmarks
Adaptiope is a domain adaptation dataset with 123 classes in the three domains synthetic, product and real life.
9 papers · 0 benchmarks
To systematically evaluate the effectiveness of our approach at accomplishing this, we designed a new benchmark, AdvBench, based on two distinct settings.
9 papers · 0 benchmarks
This dataset contains 8.9M commonsense assertions extracted by the Ascent pipeline developed at the Max Planck Institute for Informatics.
9 papers · 0 benchmarks
BIMCV-COVID19+ dataset is a large dataset with chest X-ray images CXR (CR, DX) and computed tomography (CT) imaging of COVID-19 patients along with their radiographic findings, pathologies, polymerase chain reaction (PCR), immunoglobulin G…
9 papers · 0 benchmarks
BiToD is a bilingual multi-domain dataset for end-to-end task-oriented dialogue modeling.
9 papers · 0 benchmarks
The Japanese-English business conversation corpus, namely Business Scene Dialogue corpus, was constructed in 3 steps: 1.
9 papers · 2 benchmarks
CHAOS (CHAOS - Combined (CT-MR) Healthy Abdominal Organ Segmentation)
CHAOS challenge aims the segmentation of abdominal organs (liver, kidneys and spleen) from CT and MRI data.
9 papers · 0 benchmarks
CQASUMM is a dataset for CQA (Community Question Answering) summarization, constructed from the 4.4 million Yahoo!
9 papers · 0 benchmarks
CaSiNo is a dataset of 1030 negotiation dialogues in English.
9 papers · 0 benchmarks
We provide manual annotations of 14 semantic keypoints for 100,000 car instances (sedan, suv, bus, and truck) from 53,000 images captured from 18 moving cameras at Multiple intersections in Pittsburgh, PA.
9 papers · 2 benchmarks
Chest X-ray images for pneumonia detection.
9 papers · 2 benchmarks
ClueWeb22 is the newest iteration of the ClueWeb line of datasets, provides 10 billion web pages affiliated with rich information.
9 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
9 papers · 0 benchmarks
Deep Learning Hard (DL-HARD) is an annotated dataset designed to more effectively evaluate neural ranking models on complex topics.
9 papers · 0 benchmarks
Defects4J is a collection of reproducible bugs and a supporting infrastructure with the goal of advancing software engineering research.
9 papers · 1 benchmark
E-KAR (Benchmark for Explainable Knowledge-intensive Analogical Reasoning)
The ability to recognize analogies is fundamental to human cognition.
9 papers · 0 benchmarks
The Earning Calls dataset consists of processed earning conference calls data (text and audio).
9 papers · 0 benchmarks
EgoProceL is a large-scale dataset for procedure learning.
9 papers · 0 benchmarks
FinRED is a relation extraction dataset curated from financial news and earning call transcripts containing relations from the finance domain.
9 papers · 0 benchmarks
GeoWebNews provides test/train examples and enable fine-grained Geotagging and Toponym Resolution (Geocoding).
9 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.