Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 38 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 1777–1824 of 3,998
The CHILI-100K dataset is a large-scale graph dataset (with overall >183M nodes, >1.2B edges) of nanomaterials generated from experimentally determined crystal structures.
2 papers · 8 benchmarks
The CHILI-3K dataset is a medium-scale graph dataset (with overall >6M nodes, >49M edges) of mono-metallic oxide nanomaterials generated from 12 selected crystal types.
2 papers · 8 benchmarks
CIC IoT Dataset 2022 This project aims to generate a state-of-the-art dataset for profiling, behavioural analysis, and vulnerability testing of different IoT devices with different protocols such as IEEE 802.11, Zigbee-based and Z-Wave.
2 papers · 0 benchmarks
This provides a benchmark for cyclist's orientation detection, "CIMAT-Cyclist" with bounding box based labels according to eight different classes depending on the orientation.
2 papers · 1 benchmark
CLEAR-Bias (Corpus for Linguistic Evaluation of Adversarial Robustness against Bias)
CLEAR-Bias is a benchmark dataset designed to evaluate the robustness of large language models (LLMs) against bias elicitation, particularly under adversarial conditions.
2 papers · 0 benchmarks
COFAR (Commonsense and Factual Reasoning in Image Search)
The COFAR (COmmonsense and FActual Reasoning) dataset is a collection of images and text queries specifically designed to challenge and evaluate image search models that aim to go beyond simple visual matching.
2 papers · 1 benchmark
CORE (Company Relation Extraction)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
CPPE-5 (Medical Personal Protective Equipment Dataset)
CPPE - 5 (Medical Personal Protective Equipment) is a new challenging dataset with the goal to allow the study of subordinate categorization of medical personal protective equipments, which is not possible with other popular data sets that…
2 papers · 1 benchmark
CQR (Contextual Query Rewrite)
CQR is an extension to the Stanford Dialogue Corpus.
2 papers · 0 benchmarks
CREPE is QA dataset containing a natural distribution of presupposition failures from online information-seeking forums.
2 papers · 0 benchmarks
A fundamental characteristic common to both human vision and natural language is their compositional nature.
2 papers · 1 benchmark
CUP (Context-sitUated Pun) is a dataset containing 4.5k tuples of context words and pun pairs, each labelled with whether they are compatible for composing a pun.
2 papers · 0 benchmarks
CVB (Video Dataset of Cattle Visual Behaviors)
Existing image/video datasets for cattle behavior recognition are mostly small, lack well-defined labels, or are collected in unrealistic controlled environments.
2 papers · 0 benchmarks
CWD30 (Crop Weed Dataset 30 species)
CWD30 comprises over 219,770 high-resolution images of 20 weed species and 10 crop species, encompassing various growth stages, multiple viewing angles, and environmental conditions.
2 papers · 0 benchmarks
CaFFe (CAlving Fronts and where to Find thEm)
The temporal variability in calving front positions of marine-terminating glaciers permits inference on the frontal ablation.
2 papers · 2 benchmarks
The Caltech Cars dataset consists of 126 rear-view photographs captured within parking lots.
2 papers · 1 benchmark
CamGes (Cambridge Hand Gesture Dataset)
The size of the data set is about 1GB.
2 papers · 1 benchmark
Causal Triplet is a causal representation learning benchmark featuring not only visually more complex scenes, but also two crucial desiderata commonly overlooked in previous works: 1) An actionable counterfactual setting, where only…
2 papers · 0 benchmarks
ChEMBL is a manually curated database of bioactive molecules with drug-like properties.
2 papers · 0 benchmarks
Set of landmark annotations for JSRT, Montgomery, Shenzhen and a subset of Padchest datasets
2 papers · 0 benchmarks
CholecT40 is the first endoscopic dataset introduced to enable research on fine-grained action recognition in laparoscopic surgery.
2 papers · 1 benchmark
CholecTrack20 (Multi-Perspective Multi-Class Multi-Object Tracking Dataset For Surgical Tools)
CholecTrack20 is a surgical video dataset focusing on laparoscopic cholecystectomy and designed for surgical tool tracking, featuring 20 annotated videos.
2 papers · 0 benchmarks
Description - Venue: NeurIPS 2024 D&B Spotlight - Repository: Code, Page, Data - Paper: arxiv.org/abs/2406.18522 - Point of Contact: Shenghai Yuan Citation If you find our paper and code useful in your research, please consider giving a…
2 papers · 0 benchmarks
The Circa (meaning ‘approximately’) dataset aims to help machine learning systems to solve the problem of interpreting indirect answers to polar questions.
2 papers · 0 benchmarks
CiteWorth is a a large, contextualized, rigorously cleaned labelled dataset for cite-worthiness detection built from a massive corpus of extracted plain-text scientific documents.
2 papers · 0 benchmarks
The dataset was created to address the crucial need for effective Extreme Weather Events Detection (EWED), an increasingly urgent task due to the rising frequency of such events driven by global warming.
2 papers · 0 benchmarks
This dataset is created from MIMIC-III (Medical Information Mart for Intensive Care III) and contains simulated patient admission notes.
2 papers · 4 benchmarks
A test dataset that annotated articles in 2020 following the CoNLL-2003 NER task.
2 papers · 1 benchmark
CoVaxFrames includes 113 Vaccine Hesitancy Framings found on Twitter about the COVID-19 vaccines.
2 papers · 0 benchmarks
CoVaxLies v2 includes 47 Misinformation Targets (MisTs) found on Twitter about the COVID-19 vaccines.
2 papers · 0 benchmarks
The Common Crawl corpus contains petabytes of data collected over 12 years of web crawling.
2 papers · 0 benchmarks
Comparative Question Completion is a dataset to evaluate what do large Language Models learn.
2 papers · 0 benchmarks
The appearance of the world varies dramatically not only from place to place but also from hour to hour and month to month.
2 papers · 1 benchmark
The Curiosity dataset consists of 14K dialogs (with 181K utterances) with fine-grained knowledge groundings, dialog act annotations, and other auxiliary annotation.
2 papers · 0 benchmarks
Archive of Global Tropical Cyclone Tracks Tracks from 1980 to May 2019.
2 papers · 0 benchmarks
CytoImageNet (CytoImageNet: A large-scale pretraining dataset for bioimage transfer learning)
CytoImageNet is a large-scale pretraining dataset of microscopy images (890K, 894 classes).
2 papers · 0 benchmarks
The DAPlankton dataset consists of over 110k expert-labeled plankton images.
2 papers · 0 benchmarks
Object Detection data set created from the engine DeepGTAV, which is based on the video game GTAV.
2 papers · 0 benchmarks
Object Detection data set created from the engine DeepGTAV, which is based on the video game GTAV.
2 papers · 0 benchmarks
DISL (Fueling Research with A Large Dataset of Solidity Smart Contracts)
DISL The full dataset report is available at: https://arxiv.org/abs/2403.16861 The DISL dataset features a collection of 514, 506 unique Solidity files that have been deployed to Ethereum mainnet.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
DRACO20K dataset is used for evaluating object canonicalization on methods that estimate a canonical frame from a monocular input image.
2 papers · 0 benchmarks
DRI Corpus (Dr. Inventor Multi-layer Scientific Corpus)
The Dr.
2 papers · 2 benchmarks
In this paper, we introduce a novel benchmarking framework designed specifically for evaluations of data science agents.
2 papers · 0 benchmarks
DSIOD (Driving Scenario Input Output Dataset)
This dataset contains data which enables the evaluation of metamodels and approches for targeted test case selection without setting up test environments or performing test runs.
2 papers · 0 benchmarks
DTGB (Dynamic Text-attributed Graph Benchmark)
We introduce Dynamic Text-attributed Graph Benchmark (DTGB), a collection of large-scale, time-evolving graphs from diverse domains, with nodes and edges enriched by dynamically changing text attributes and categories.
2 papers · 0 benchmarks
DUC 2007 (Document Understanding Conferences)
There is currently much interest and activity aimed at building powerful multi-purpose information systems.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.