Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 48 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 2257–2304 of 3,998
AutoFR Dataset is broken down by each site that we crawl within a zip file.
1 paper · 0 benchmarks
This is a benchmark for neural paraphrase detection, to differentiate between original and machine-generated content.
1 paper · 0 benchmarks
Dataset Overview: 998 images and 4,208 annotations focusing on interaction with in-vehicle infotainment (IVI) systems.
1 paper · 0 benchmarks
For more details see https://huggingface.co/datasets/jpwahle/autoregressive-paraphrase-dataset
1 paper · 0 benchmarks
AuxAD is a a distantly supervised dataset for acronym disambiguation.
1 paper · 0 benchmarks
AuxAI is a distantly supervised dataset for acronym identification.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
AviationQA is introduced in the paper titled- There is No Big Brother or Small Brother: Knowledge Infusion in Language Models for Link Prediction and Question Answering The paper is accepted in the main conference of ICON 2022.
1 paper · 1 benchmark
BAH (Behavioural Ambivalence/Hesitancy)
Recognizing complex emotions linked to ambivalence and hesitancy (A/H) can play a critical role in the personalization and effectiveness of digital behaviour change interventions.
1 paper · 0 benchmarks
BASIR (BASIR_Budget_Assisted_Sectoral_Impact_Ranking)
Government fiscal policies, particularly annual union budgets, exert significant influence on financial markets.
1 paper · 0 benchmarks
This dataset is for evaluating the task of Black-box Multi-agent Integration which focuses on combining the capabilities of multiple black-box conversational agents at scale.
1 paper · 1 benchmark
BCSD (Bank Check Segmentation Dataset)
The dataset consists of images of 158 filled out bank checks containing various complex backgrounds, and handwritten text and signatures in the respective fields, along with both pixel-level and patch-level segmentation masks for the…
1 paper · 0 benchmarks
BD-TypoSAT (Building Damage Typology Satellite Dataset)
On Sunday, August 29, 2021, Hurricane Ida struck parts of Louisiana and Mississippi with wind gusts reaching up to 172 mph, leaving more than a million customers without electricity, including the entire New Orleans area.
1 paper · 0 benchmarks
BDD-QA is distinguished by its encompassing range of traffic actions, crafted to rigorously evaluate a model's decision-making abilities in traffic scenario.
1 paper · 0 benchmarks
BEAR-probe (Benchmark for Evaluating Associative Reasoning)
The BEAR dataset and its larger version, BEARbig, are benchmarks for evaluating common factual knowledge contained in language models.
1 paper · 0 benchmarks
BERSt (Basic Emotion Random phrase Shouts)
BERSt Dataset We release the BERSt Dataset for various speech recognition tasks including Automatic Speech Recognition (ASR) and Speech Emotion Recogniton (SER) Overview 4526 single phrase recordings (~3.75h) 98 professional actors 19…
1 paper · 1 benchmark
BFN (Backdoored Face-Networks Dataset)
This database is a database of backdoored neural networks intended for face recognition.
1 paper · 0 benchmarks
BIDCD (Bosch Industrial Depth Completion Dataset)
Bosch Industrial Depth Completion Dataset (BIDCD) is an RGBD dataset for of static table-top scenes with industrial objects.
1 paper · 0 benchmarks
This dataset is a BIDS-compatible version of the CHB-MIT Scalp EEG Database.
1 paper · 0 benchmarks
This dataset is a BIDS compatible version of the Siena Scalp EEG Database.
1 paper · 0 benchmarks
BIOSED-ACPD: BIOacoustic Sound Event Detection - Adaptive Change Point Detection dataset Description.
1 paper · 0 benchmarks
The BIRDeep Audio Annotations dataset is a collection of bird vocalizations from Doñana National Park, Spain.
1 paper · 0 benchmarks
BLANCA (Benchmarks for LANguage models on Coding Artifacts) is a collection of benchmarks that assess code understanding based on tasks such as predicting the best answer to a question in a forum post, finding related forum posts, or…
1 paper · 0 benchmarks
BLM-17m is a labeled dataset for topic detection that contains 17 million tweets.
1 paper · 0 benchmarks
BLN600 (BLN600: A Parallel Corpus of Machine/Human Transcribed Nineteenth Century Newspaper Texts)
A publicly available corpus of nineteenth-century newspaper text focused on crime in London, derived from the Gale British Library Newspapers corpus parts 1 and 2.
1 paper · 0 benchmarks
BLP (Blackout Poetry Dataset)
A blackout poetry dataset constructed from publicly available short stories and large poems.
1 paper · 1 benchmark
BP^C (A Benchmark Dataset for Causal Business Process Reasoning)
Large Language Models (LLMs) are increasingly used for boosting organizational efficiency and automating tasks.
1 paper · 0 benchmarks
BPersona-chat is an evaluation dataset based on the English multiturn chat corpus Persona-chat and the Japanese multiturn chat corpus JPersona-chat.
1 paper · 0 benchmarks
The BWB corpus consists of Chinese novels translated by experts into English, and the annotated test set is designed to probe the ability of machine translation systems to model various discourse phenomena.
1 paper · 0 benchmarks
This is a link to the source code of the Baking-Large domain introduced in the paper.
1 paper · 0 benchmarks
The dataset consists of images of bananas and apples.
1 paper · 0 benchmarks
A Bilingual Dataset for Bangla and English Voice Commands Colloquial Bangla has adopted many English words due to colonial influence.
1 paper · 1 benchmark
Millions of people around the world have low or no vision.
1 paper · 0 benchmarks
The dataset identifies the shortcomings of existing benchmarks in evaluating the problem of compositional generalization, which underscores the need for the development of datasets tailored to assess compositional generalization in open…
1 paper · 1 benchmark
battery surface temperature from 16 sensors.
1 paper · 0 benchmarks
Forty prismatic lithium-ion pouch cells were built at the University of Michigan Battery Laboratory.
1 paper · 0 benchmarks
In this dataset two robots, Baxter and UR5, perform 8 behaviors (look, grasp, pick, hold, shake, lower, drop, and push) on 95 objects that vary by 5 color (blue, green, red, white, and yellow), 6 contents (wooden button, plastic dices,…
1 paper · 0 benchmarks
Bc8BioRED is built upon BioRED 2022 with the addition of directionality annotations.
1 paper · 1 benchmark
Beemo (Benchmark of expert-edited machine-generated outputs)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
BenBench is designed to benchmark the potential for data leakage in benchmark datasets, which can lead to biased and inequitable comparisons.
1 paper · 0 benchmarks
BestRev (Understanding Peer Review of Software Engineering Papers)
Survey instrument, analysis code, and anonymized responses for the paper on review practices in SE.
1 paper · 0 benchmarks
The dataset covers three types of medical interactions in both English and Arabic: - Multiple-choice question answering (MCQA), focusing on specialized medical knowledge.
1 paper · 0 benchmarks
Bianet is a parallel news corpus in Turkish, Kurdish and English It contains 3,214 Turkish articles with their sentence-aligned Kurdish or English translations from the Bianet online newspaper.
1 paper · 0 benchmarks
BioDrone is the first bionic drone-based single object tracking benchmark, it features videos captured from a flapping-wing UAV system with a major camera shake due to its aerodynamics.
1 paper · 0 benchmarks
BioFuelQR is a dataset consisting of complex reasoning questions related to catalyst discovery in biofuels.
1 paper · 0 benchmarks
The measurement data VNA20220722232002XETSreduced.mat includes a data matrix 𝐑 acquired with a synthetic aperture measurement testbed described in [2] and [3].
1 paper · 0 benchmarks
It includes 227 impactful events in Bitcoin history that shook the global markets.
1 paper · 0 benchmarks
📚 BlendNet The dataset contains $12k$ samples.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.