Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 18 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 817–864 of 3,998
Housekeep a benchmark to evaluate common sense reasoning in the home for embodied AI.
11 papers · 0 benchmarks
ImageCoDe (Image Retrieval from Contextual Descriptions)
Given 10 minimally contrastive (highly similar) images and a complex description for one of them, the task is to retrieve the correct image.
11 papers · 1 benchmark
LOL-v2-synthetic (From Fidelity to Perceptual Quality: A Semi-Supervised Approach for Low-Light Image Enhancement)
From Fidelity to Perceptual Quality: A Semi-Supervised Approach for Low-Light Image Enhancement
11 papers · 1 benchmark
The MAGE dataset provides a large set of generated texts using 27 LLMs from seven different groups: OpenAI GPT, LLaMA, GLM130B, FLAN-T5, OPT, BigScience, and EleutherAI.
11 papers · 1 benchmark
Multi-Modal Reading (MMR) Benchmark includes 550 annotated question-answer pairs across 11 distinct tasks involving texts, fonts, visual elements, bounding boxes, spatial relations, and grounding, with carefully designed evaluation metrics.
11 papers · 1 benchmark
We scrape data from GooBix, which contains 156 games of 5 × 5 mini crosswords.
11 papers · 0 benchmarks
MultiEURLEX is a multilingual dataset for topic classification of legal documents.
11 papers · 0 benchmarks
The NewSHead dataset contains 369,940 English stories with 932,571 unique URLs, among which there are 359,940 stories for training, 5,000 for validation, and 5,000 for testing, respectively.
11 papers · 1 benchmark
Ohsumed includes medical abstracts from the MeSH categories of the year 1991.
11 papers · 2 benchmarks
OpenLane-V2 is the world's first perception and reasoning benchmark for scene structure in autonomous driving.
11 papers · 2 benchmarks
P-DukeMTMC-reID is a modified version based on DukeMTMC-reID dataset.
11 papers · 1 benchmark
PATS (Pose Audio Transcript Style)
PATS dataset consists of a diverse and large amount of aligned pose, audio and transcripts.
11 papers · 0 benchmarks
ParCorFull (Parallel Corpus Annotated with Full Coreference)
ParCorFull is a parallel corpus annotated with full coreference chains that has been created to address an important problem that machine translation and other multilingual natural language processing (NLP) technologies face -- translation…
11 papers · 0 benchmarks
A large-scale collection of visually-grounded, task-oriented dialogues in English designed to investigate shared dialogue history accumulating during conversation.
11 papers · 0 benchmarks
ReCAM (SemEval-2021 Task 4: Reading Comprehension of Abstract Meaning)
Tasks Our shared task has three subtasks.
11 papers · 1 benchmark
SMAC-Exp (StarCraft Multi-Agent Exploration Challenge)
The StarCraft Multi-Agent Challenges+ requires agents to learn completion of multi-stage tasks and usage of environmental factors without precise reward functions.
11 papers · 2 benchmarks
SOMOS (The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis)
The SOMOS dataset is a large-scale mean opinion scores (MOS) dataset consisting of solely neural text-to-speech (TTS) samples.
11 papers · 0 benchmarks
StylePTB is a fine-grained text style transfer benchmark.
11 papers · 0 benchmarks
Synbols is a dataset generator designed for probing the behavior of learning algorithms.
11 papers · 0 benchmarks
The first large demoire dataset.
11 papers · 1 benchmark
TRIPOD (TuRnIng POint Dataset)
TRIPOD contains screenplays and plot synopses with turning point (TP) annotations for 99 movies.
11 papers · 0 benchmarks
Talk The Walk is a large-scale dialogue dataset grounded in action and perception.
11 papers · 0 benchmarks
TaxiNLI is a dataset collected based on the principles and categorizations of the aforementioned taxonomy.
11 papers · 0 benchmarks
TimeDial presents a crowdsourced English challenge set, for temporal commonsense reasoning, formulated as a multiple choice cloze task with around 1.5k carefully curated dialogs.
11 papers · 0 benchmarks
UDIS-D (Unsupervised Deep Image Stitching Dataset)
UDIS-D is a large image dataset for image stitching or image registration.
11 papers · 0 benchmarks
VoxForge is an open speech dataset that was set up to collect transcribed speech for use with Free and Open Source Speech Recognition Engines (on Linux, Windows and Mac).
11 papers · 9 benchmarks
Robust detection and tracking of objects is crucial for the deployment of autonomous vehicle technology.
11 papers · 2 benchmarks
nvBench is a large-scale NL2VIS (natural languagge to visualisations) benchmark, containing 25,750 (NL, VIS) pairs from 750 tables over 105 domains, synthesized from (NL, SQL) benchmarks to support cross-domain NLPVIS (Natural Language…
11 papers · 0 benchmarks
The AI City Challenge, hosted at CVPR 2024, focuses on harnessing AI to enhance operational efficiency in physical settings such as retail and warehouse environments, and Intelligent Traffic Systems (ITS).
10 papers · 1 benchmark
ABCD (Action-Based Conversations Dataset)
10 papers · 1 benchmark
AnnoMI: A Dataset of Expert-Annotated Counselling Dialogues Dataset Introduction Research on natural language processing approaches to analysing counselling dialogues has seen substantial development in recent years, but access to this…
10 papers · 0 benchmarks
BiRD (Bigram Relatedness Dataset)
Bigram Relatedness Dataset (BiRD) is a large, fine-grained, bigram relatedness dataset, using a comparative annotation technique called Best Worst Scaling.
10 papers · 0 benchmarks
ChatHaruhi (ChatHaruhi: Reviving Anime Character in Reality via Large Language Model)
ChatHaruhi is a dataset covering 32 Chinese / English TV / anime characters with over 54k simulated dialogues.
10 papers · 0 benchmarks
CriticBench is a comprehensive benchmark designed to assess the abilities of Large Language Models (LLMs) to critique and rectify their reasoning across various tasks.
10 papers · 0 benchmarks
The CropAndWeed dataset is focused on the fine-grained identification of 74 relevant crop and weed species with a strong emphasis on data variability.
10 papers · 0 benchmarks
DDPM (Deception Detection and Physiological Monitoring)
The Deception Detection and Physiological Monitoring (DDPM) dataset captures an interview scenario in which the interviewee attempts to deceive the interviewer on selected responses.
10 papers · 0 benchmarks
Although deep face recognition has achieved impressive results in recent years, there is increasing controversy regarding racial and gender bias of the models, questioning their trustworthiness and deployment into sensitive scenarios.
10 papers · 0 benchmarks
The Discovery datasets consists of adjacent sentence pairs (s1,s2) with a discourse marker (y) that occurred at the beginning of s2.
10 papers · 1 benchmark
This is the dataset for the 2020 Duolingo shared task on Simultaneous Translation And Paraphrase for Language Education (STAPLE).
10 papers · 0 benchmarks
Satellite images are snapshots of the Earth surface.
10 papers · 4 benchmarks
EmoWOZ is the first large-scale open-source dataset for emotion recognition in task-oriented dialogues.
10 papers · 2 benchmarks
FM-IQA (Freestyle Multilingual Image Question Answering)
FM-IQA is a question-answering dataset containing over 150,000 images and 310,000 freestyle Chinese question-answer pairs and their English translations.
10 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
10 papers · 0 benchmarks
Fig-QA consists of 10256 examples of human-written creative metaphors that are paired as a Winograd schema.
10 papers · 0 benchmarks
The HO-3D v3 is the version 3 of the HO-3D dataset with more accurate hand-object poses.
10 papers · 1 benchmark
ImageNet-VidVRD dataset contains 1,000 videos selected from ILVSRC2016-VID dataset based on whether the video contains clear visual relations.
10 papers · 2 benchmarks
Dataset Introduction In this work, we introduce the In-Diagram Logic (InDL) dataset, an innovative resource crafted to rigorously evaluate the logic interpretation abilities of deep learning models.
10 papers · 1 benchmark
LectureBank Dataset is a manually collected dataset of lecture slides.
10 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.