Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 51 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 2401–2448 of 3,998

Scene change detection (SCD) dataset tailored for generalizable SCD algorithm.
1 paper · 1 benchmark
Taking Advice from ChatGPT is a laboratory study of how student participants incorporate advice generated by ChatGPT.
1 paper · 0 benchmarks
Dataset Overview vanilla.csv: Represents the interactions without specific role-play instructions.
1 paper · 0 benchmarks
The dataset contains two few-shot chemical fine-grained entity extraction datasets, based on human-annotated ChemNER+ and CHEMET.
1 paper · 0 benchmarks
ChessReD (Chess Recognition Dataset)
The Chess Recognition Dataset (ChessReD) comprises a diverse collection of images of chess formations captured using smartphone cameras; a sensor choice made to ensure real-world applicability.
1 paper · 0 benchmarks
ChessReD2K (Chess Recognition Dataset 2K)
The Chess Recognition Dataset 2K (ChessReD2K) comprises a diverse collection of images of chess formations captured using smartphone cameras; a sensor choice made to ensure real-world applicability.
1 paper · 0 benchmarks
ChildCIdb (ChildCIdbv1)
A large-scale, first-of-its-kind database aimed at generating a better understanding of the way children interact with mobile devices during their development process.
1 paper · 0 benchmarks
ChinaOpen is a new video dataset targeted at open-world multimodal learning, with raw data gathered from Bilibili, a popular Chinese video-sharing website.
1 paper · 1 benchmark
Description - Venue: NeurIPS 2024 D&B Spotlight - Repository: Code, Page, Data - Paper: arxiv.org/abs/2406.18522 - Point of Contact: Shenghai Yuan Citation If you find our paper and code useful in your research, please consider giving a…
1 paper · 0 benchmarks
This dataset contains a two-column CSV file, where the first column ("ValidcitingDOI") contains the DOI of a citing entity retrieved in Crossref, while the second column ("InvalidcitedDOI") contains the invalid DOI of a cited entity…
1 paper · 0 benchmarks
- An evaluation test bed for assessing the robustness of sentence embedding models against user-informed misinformation edits.
1 paper · 0 benchmarks
Clarkson Fingerprint Generator consists of a dataset of 50K synthetically generated fingerprints.
1 paper · 0 benchmarks
Clickable heat-map visualizations of the experiments run to quantify the Classic ECN AQM problem and to evaluate the success of the Classic AQM Detection and Fall-back algorithm.
1 paper · 0 benchmarks
Lang-8 Preprocessed Dataset (for GED): - Dataset: Lang-8, a publicly available dataset containing user-generated content, primarily from second-language learners, focused on writing errors.
1 paper · 0 benchmarks
Clickbait PDFs (From Attachments to SEO: Click Here to Learn More about Clickbait PDFs!)
The paper presents a study of Clickbait PDFs, which are PDF documents leading to various attacks on the Web.
1 paper · 0 benchmarks
ClustMe and ClustML data S1 and S2 (Gaussian Mixture human-labeled data for clustering design and evaluation)
1 paper · 0 benchmarks
CoNECo (Complex Named Entity Corpus)
Complex Named Entity Corpus (CoNECo) is an annotated corpus for NER and NEN of protein-containing complexes.
1 paper · 0 benchmarks
CoSQA+ (CoSQA_Plus)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Dataset used in research submitted to ICPC ERA 2022
1 paper · 0 benchmarks
Code and Data for Replication of "Microsimulation Estimates of Decision Uncertainty and Value of Information Are Biased but Consistent" This is the full data set for replication of all results in the paper along with the R code for doing…
1 paper · 0 benchmarks
CodeInstruct (InstructCoder, CodeInstruct)
InstructCoder is the first dataset designed to adapt LLMs for general code editing.
1 paper · 0 benchmarks
ColorSVG-100K contains: - 100K samples - 500 categories Project website
1 paper · 0 benchmarks
A large dataset of color names and their respective RGB values stores in CSV.
1 paper · 1 benchmark
Probes to evaluate commonsense in language models.
1 paper · 0 benchmarks
CompMix-IR Dataset Overview: Characteristics: CompMix-IR is a heterogeneous knowledge retrieval benchmark dataset, featuring four knowledge types (text, knowledge graphs, tables, and infoboxes), 9,400+ QA pairs, and a corpus of 10 million…
1 paper · 0 benchmarks
The 50-ha plot at Barro Colorado Island was initially demarcated and fully censused in 1982, and has been fully censused 7 times since, every 5 years from 1985 through 2015.
1 paper · 0 benchmarks
In everyday language processing, sentence context affects how readers and listeners process upcoming words.
1 paper · 0 benchmarks
The Composed Quora dataset consists of questions extracted from Quora that are grouped together if they are asking the same thing.
1 paper · 0 benchmarks
Computer Vision Arxiv Figures dataset consists of 88,645 images that more closely resemble the structure of our visual prompts.
1 paper · 0 benchmarks
ConCon Dataset (Continually Confounded Dataset)
ConCon: Continually Confounded Dataset is a confounded visual dataset for continual learning.
1 paper · 0 benchmarks
ConQA (Conceptual Query Answering)
ConQA is a dataset created using the intersection between VisualGenome and MS-COCO.
1 paper · 2 benchmarks
ConSLAM (Construction Dataset for SLAM)
ConSLAM is a real-world dataset collected periodically on a construction site to measure the accuracy of mobile scanners' SLAM algorithms.
1 paper · 0 benchmarks
Concept-1K contains 1023 novel concepts from six domains, including economy, culture, science and technology, environment, education, and health and medical.
1 paper · 0 benchmarks
Concrete is the most important material in civil engineering.
1 paper · 1 benchmark
Description - Repository: Code, Page, Data - Paper: arxiv.org/abs/2411.17440 - Point of Contact: Shenghai Yuan Citation If you find our paper and code useful in your research, please consider giving a star and citation.
1 paper · 0 benchmarks
Enumerate–Conjecture–Prove: Formally Solving Answer-Construction Problem in Math Competitions We release the ConstructiveBench dataset as part of our Enumerate–Conjecture–Prove (ECP) paper.
1 paper · 0 benchmarks
This is a revised and extended second version of a Contextualised Polyseme Word Sense Dataset.
1 paper · 0 benchmarks
Corpus of controversial news articles extracted from Twitter.
1 paper · 0 benchmarks
ConvSumX is a cross-lingual conversation summarization benchmark, through a new annotation schema that explicitly considers source input context.
1 paper · 0 benchmarks
This dataset contains 12,500 meter images acquired in the field by the employees of the Energy Company of Paraná (Copel), which directly serves more than 4 million consuming units, across 395 cities and 1,113 locations (i.e., districts,…
1 paper · 1 benchmark
CoronaVis is a dataset of tweets related to coronavirus.
1 paper · 0 benchmarks
Dataset of Stack Overflow questions about Terraform with cost-related keywords.
1 paper · 0 benchmarks
Dataset of commit messages and issues containing evidence of cost awareness.
1 paper · 0 benchmarks
Using Council Data Project infrastructures (https://councildataproject.org), we assemble longitudinal municipal council meeting transcript data.
1 paper · 0 benchmarks
Counting Probe (Counting Probe based on Visual7W)
Probing cross-modal capabilities of Vision & Language models with a counting task.
1 paper · 0 benchmarks
Creative Visual Storytelling Anthology (ARL Creative Visual Storytelling Anthology)
The Creative Visual Storytelling Anthology is a collection of 100 author responses to an improved creative visual storytelling exercise over a sequence of three images.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.