Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 26 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 1201–1248 of 3,998

Click to add a brief description of the dataset (Markdown and LaTeX enabled).
5 papers · 0 benchmarks
The Epinions dataset is trust network dataset.
5 papers · 1 benchmark
FLD (Formal Logic Deduction)
A deductive reasoning benchmark based on formal logic theory.
5 papers · 0 benchmarks
Contains 8k flickr Images with captions.
5 papers · 2 benchmarks
GOF (Gyroscope Optical Flow)
Optical Flow in challenging scenes with gyroscope readings!
5 papers · 0 benchmarks
The first AI-generated video detection datasets.
5 papers · 0 benchmarks
Geolife (Geolife GPS Trajectory Dataset)
This is a GPS trajectory dataset collected in (Microsoft Research Asia) GeoLife project by 182 users in a period of over three years (from April 2007 to August 2012).
5 papers · 0 benchmarks
Grep-BiasIR (Gender Representation-Bias for Information Retrieval)
Grep-BiasIR is a novel thoroughly-audited dataset which aim to facilitate the studies of gender bias in the retrieved results of IR systems.
5 papers · 0 benchmarks
Benchmark for de novo molecular design
5 papers · 0 benchmarks
HaDes is a token-level, reference-free hallucination detection dataset named HAllucination DEtection dataSet.
5 papers · 0 benchmarks
HaN-Seg (The Head and Neck Organ-at-Risk CT & MR Segmentation Challenge)
Cancer in the region of the head and neck (HaN) is one of the most prominent cancers, for which radiotherapy represents an important treatment modality that aims to deliver a high radiation dose to the targeted cancerous cells while…
5 papers · 0 benchmarks
The Headlines dataset for sarcasm detection is collected from two news website.
5 papers · 0 benchmarks
We introduce HourVideo, a benchmark dataset for hour-long video-language understanding.
5 papers · 0 benchmarks
HowMany-Qa is a object counting dataset.
5 papers · 1 benchmark
A dataset of 69,270,581 video clip, question and answer triplets (v, q, a).
5 papers · 0 benchmarks
Hummingbird is a dataset to examine stylistic lexical cues from human perception and BERT used to characterize their discrepancy.
5 papers · 0 benchmarks
IAM(line-level) (Line-level Handwritten Text Recognition on IAM)
The IAM database contains 13,353 images of handwritten lines of text created by 657 writers.
5 papers · 1 benchmark
Illness-dataset (Illness multi-domain textual dataset)
A dataset for evaluating text classification, domain adaptation, and active learning models.
5 papers · 0 benchmarks
ImageNet-Hard is a new benchmark that comprises 10,980 images collected from various existing ImageNet-scale benchmarks (ImageNet, ImageNet-V2, ImageNet-Sketch, ImageNet-C, ImageNet-R, ImageNet-ReaL, ImageNet-A, and ObjectNet).
5 papers · 1 benchmark
The ImplicitQA dataset was introduced in the paper ImplicitQA: Going beyond frames towards Implicit Video Reasoning.
5 papers · 1 benchmark
InfiniBench (InfiniBench: A Comprehensive Benchmark for Large Multimodal Models in Very Long Video Understanding)
We introduce InfiniBench a comprehensive benchmark for very long video understanding, which presents 1) The longest video duration, averaging 76.34 minutes; 2) The largest number of question-answer pairs, 108.2K; 3) Diversity in questions…
5 papers · 0 benchmarks
We collect, organize and open-source the large-scale multimodal instruction dataset, Infinity-MM, consisting of tens of millions of samples.
5 papers · 0 benchmarks
InsPLAD (Inspection Power Line Asset Dataset)
InsPLAD is a Dataset for Power Line Asset Inspection containing 10,607 high-resolution Unmanned Aerial Vehicles colour images.
5 papers · 1 benchmark
Instructional-DT (Instr-DT) (Instructional Discourse Treebank)
This discourse treebank includes annotated instructional texts originally assembled at the Information Technology Research Institute, University of Brighton.
5 papers · 1 benchmark
IoT devices captures - Samuel Marchal (Creator) Description This dataset represents the traffic emitted during the setup of 31 smart home IoT devices of 27 different types (4 types are represented by 2 devices each).
5 papers · 0 benchmarks
KETOD (Knowledge-Enriched Task-Oriented Dialogue)
KETOD (Knowledge-Enriched Task-Oriented Dialogue) is a dataset containing system responses designed for enriching task-oriented dialogues with chit-chat based on relevant entity knowledge.
5 papers · 0 benchmarks
KiloGram is a resource for studying abstract visual reasoning in humans and machines.
5 papers · 0 benchmarks
Kvasir-Sessile dataset (Sessile polyps from Kvasir-SEG)
The Kvasir-SEG dataset includes 196 polyps smaller than 10 mm classified as Paris class 1 sessile or Paris class IIa.
5 papers · 0 benchmarks
LUDB (Lobachevsky University Electrocardiography Database)
Abstract Lobachevsky University Electrocardiography Database (LUDB) is an ECG signal database with marked boundaries and peaks of P, T waves and QRS complexes.
5 papers · 1 benchmark
LabPics (LabPics Dataset for computer vision for autonomous chemistry labs and medical labs)
LabPics Chemistry Dataset Dataset for computer vision for materials segmentation and classification in chemistry labs, medical labs, and any setting where materials are handled inside containers.
5 papers · 0 benchmarks
We propose a novel long-context benchmark, 🐉 Loong, aligning with realistic scenarios through extended multi-document question answering (QA).
5 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
5 papers · 0 benchmarks
MarKG (Multimodal analogical reasoning Knowledge Graph)
The MarKG dataset has 11,292 entities, 192 relations and 76,424 images, including 2,063 analogy entities and 27 analogy relations.
5 papers · 0 benchmarks
MatSynth MatSynth is a Physically Based Rendering (PBR) materials dataset designed for modern AI applications.
5 papers · 0 benchmarks
MediaSpeech is a media speech dataset (you might have guessed this) built with the purpose of testing Automated Speech Recognition (ASR) systems performance.
5 papers · 1 benchmark
MindCraft is a fine-grained dataset of collaborative tasks performed by pairs of human subjects in the 3D virtual blocks world of Minecraft.
5 papers · 0 benchmarks
MoA (MoA_Long_ModelQA)
This is the dataset used by the automatic sparse attention compression method MoA.
5 papers · 0 benchmarks
MuMiN is a misinformation graph dataset containing rich social media data (tweets, replies, users, images, articles, hashtags), spanning 21 million tweets belonging to 26 thousand Twitter threads, each of which have been semantically…
5 papers · 0 benchmarks
NELA-GT-2019 is an updated version of the NELA-GT-2018 dataset.
5 papers · 0 benchmarks
NFCorpus is a full-text English retrieval data set for Medical Information Retrieval.
5 papers · 1 benchmark
NISP (NITK-IISc Multilingual Multi-accent Speaker Profiling)
This dataset contains speech recordings along with speaker physical parameters (height, weight, shoulder size, age ) as well as regional information and linguistic information.
5 papers · 0 benchmarks
Netzschleuder (network catalogue, repository and centrifuge)
This is a catalogue and repository of network datasets with the aim of aiding scientific research.
5 papers · 0 benchmarks
Synthetically Generated Night-time Weather Degraded Database
5 papers · 1 benchmark
OoDIS (Anomaly Instance Segmentation Benchmark)
OoDIS is a benchmark dataset for anomaly instance segmentation, crucial for autonomous vehicle safety.
5 papers · 2 benchmarks
Open6DOR V2 (Benchmarking Open-instruction 6-DoF Object Rearrangement and A VLM-based Approach)
We introduce a challenging and comprehensive benchmark for open-instruction 6-DoF object rearrangement tasks, termed Open6DOR.
5 papers · 1 benchmark
PEN (Problems with Explanations for Numbers)
Provided explanations on the existing three benchmark datasets on solving algebraic word problems: ALG514, DRAW-1K, MAWPS
5 papers · 1 benchmark
PGDP5K (Plane Geometry Diagram Parsing Dataset)
PGDP5K is a dataset consisting of 5000 diagram samples composed of 16 shapes, covering 5 positional relations, 22 symbol types and 6 text types, labeled with more fine-grained annotations at primitive level, including primitive classes,…
5 papers · 1 benchmark
Year after year, the demand for ever-better smartphone photos continues to grow, in particular in the domain of portrait photography.
5 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.