Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 77 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 3649–3696 of 3,998
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
These datasets, ComCo and SimCo, designed for evaluating multi-object representation in Vision-Language Models (VLMs).
1 paper · 0 benchmarks
The simply-CLEVR dataset aims to provide a benchmark dataset that can be used for transparent quantitative evaluation of explanation methods (aka heatmaps/XAI methods).
1 paper · 0 benchmarks
This dataset is a patched version of The Taste & Affect Music Database by D.
1 paper · 0 benchmarks
The datasets of "Time Interval-enhanced Graph Neural Network for Shared-account Cross-domain Sequential Recommendation" (TNNLs 2022)
1 paper · 0 benchmarks
This deposit contains benchmark code, data and results to assess the Python software time-agnostic-library v4.3.0.
1 paper · 0 benchmarks
topex-printer is a dataset containing 102 machine parts of a label printing machine.
1 paper · 0 benchmarks
Dataset based on Twitter usernames of American politicians.
1 paper · 0 benchmarks
Microscopy is a cornerstone of biomedical research, enabling detailed study of biological structures at multiple scales.
1 paper · 0 benchmarks
The uniD dataset is an innovative collection of naturalistic road user trajectories, captured within the RWTH Aachen University campus using drone technology to address common challenges such as occlusions found in traditional traffic data…
1 paper · 0 benchmarks
This dataset contains the ground truth for urban changes occurred in Mariupol, Ukraine for the time frame 2017-2020.
1 paper · 0 benchmarks
LLM-Based Vulnerability Classification in Police Narratives This repository contains datasets used in our research on applying large language models (LLMs) to identify indicators of vulnerability in police incident narratives.
1 paper · 0 benchmarks
VQA NLE synthetic dataset, made with LLaVA-1.5 using features from GQA dataset.
1 paper · 0 benchmarks
Dataset contains about 48K contracts which are open source on Etherscan.
1 paper · 0 benchmarks
The dataset consists of 53,189 wikiHow articles across various categories of everyday tasks, 155,265 methods, and 772,294 steps with corresponding images.
1 paper · 1 benchmark
Here I provided the datasets I used for this analysis.
1 paper · 0 benchmarks
This is a stance detection dataset in the Zulu language.
1 paper · 0 benchmarks
ABODA (Abandoned Object Dataset)
ABandoned Objects DAtaset (ABODA) is a new public dataset for abandoned object detection.
0 papers · 0 benchmarks
ADFI (Anomaly Detection Datasets for Visual Inspection)
ADFI Dataset is an image dataset for anomaly detection methods with a focus on industrial inspection.
0 papers · 0 benchmarks
ALFI (Annotations for Label-Free Images)
ALFI (Annotations for Label-Free Images) is a dataset of images and annotations for label-free microscopy imaging.
0 papers · 0 benchmarks
This dataset is described in the ALTA 2022 Shared Task and associated CodaLab competition.
0 papers · 0 benchmarks
ALTA 2023 Shared Task (Discriminate between human-authored and synthetic text generated by Large Language Models (LLMs))
This dataset is described in the ALTA 2023 Shared Task and associated CodaLab competition.
0 papers · 0 benchmarks
ARF (Artificial Relationships in Fiction)
Artificial Relationships in Fiction Dataset Description Artificial Relationships in Fiction (ARF) is a synthetically annotated dataset for Relation Extraction (RE) in fiction, created from a curated selection of literary texts sourced from…
0 papers · 0 benchmarks
This open-source dataset consists of 5.04 hours of transcribed English conversational speech beyond telephony, where 13 conversations were contained.
0 papers · 0 benchmarks
Affective Text (Test Corpus of SemEval 2007) by Carlo Strapparava & Rada Mihalcea.
0 papers · 0 benchmarks
AlexMI (Alex Motor Imagery dataset)
Alex Motor Imagery dataset.
0 papers · 0 benchmarks
AntM2C (Ant-Group Multi-Scenario Multi-Modal CTR dataset)
We release a large-scale Multi-Scenario Multi-Modal CTR dataset named AntM2C, built from real industrial data from Alipay.
0 papers · 0 benchmarks
Dataset contains images with apples infected by scab.
0 papers · 0 benchmarks
Dataset contains images with apple leaves infected by scab.
0 papers · 0 benchmarks
Global epidemics, like COVID-19, have substantial impacts on almost all countries in multiple aspects, such as economy, hospitalization, lifestyle, etc1, 2.
0 papers · 0 benchmarks
ArcBench is a logically challenging dataset of 158 English question–answer pairs, derived from the RoR-Bench benchmark.
0 papers · 0 benchmarks
The Arena-Hard benchmark is a high-quality benchmarking tool for Language Learning Models (LLMs) developed by LMSYS Org¹.
0 papers · 0 benchmarks
Data files with the information required to replicate all the experiments reported in the paper: Linares López, Carlos; Herman, Ian, 2024.
0 papers · 0 benchmarks
This dataset is comprised of the dynamic analysis reports generated by CAPEv2, from both malware and goodware.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.