Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 66 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 3121–3168 of 3,998

ProofNetVerif is an evaluation benchmark comprising 3,752 entries, each including an informal mathematical statement, its reference formalization, a predicted formalization, and a binary label indicating semantic equivalence.
1 paper · 0 benchmarks
OOD split of the Mol-Instructions Dataset about Protein Annotation.
1 paper · 0 benchmarks
This is a large-scale court judgment dataset, where each judgment is a summary of the case description with a patternized style.
1 paper · 0 benchmarks
PubChem18 (PubChem 2018)
A.2.1 AN OPEN, LARGE-SCALE DATASET FOR ZERO-SHOT DRUG DISCOVERY DERIVED FROM PUBCHEM We constructed a large public dataset extracted from PubChem (Kim et al., 2019; Preuer et al., 2018), an open chemistry database, and the largest…
1 paper · 0 benchmarks
he goal of this project is to use high throughput screening approaches to identify and develop novel, highly selective small molecule allosteric modulators of the D1 DAR for use as in vitro and in vivo pharmacological tools and in…
1 paper · 0 benchmarks
This dataset gathers 14,857 entities, 133 relations, and entities corresponding tokenized text from PubMed.
1 paper · 0 benchmarks
This dataset gathers three types of pairs: Title-to-Abstract (Training: 22,811/Development: 2095/Test: 2095), Abstract-to-Conclusion and Future work (Training: 22,811/Development: 2095/Test: 2095), Conclusion and Future work-to-Title…
1 paper · 0 benchmarks
This dataset contains a probabilistic sample of ~2.4 million PubMed abstracts, enriched with precomputed dense embeddings (title + abstract), from the ncbi/MedCPT-Article-Encoder model.
1 paper · 0 benchmarks
PubMedQA-MetaGen: Metadata-Enriched PubMedQA Corpus Dataset Summary PubMedQA-MetaGen is a metadata-enriched version of the PubMedQA biomedical question-answering dataset, created using the MetaGenBlendedRAG enrichment pipeline.
1 paper · 2 benchmarks
A benchmark for Python library migration.
1 paper · 0 benchmarks
Pylon Benchmark (Pylon Table Union Search Benchmark)
We create a new dataset from GitTables, a data lake of 1.7M tables extracted from CSV files on GitHub.
1 paper · 0 benchmarks
QASports (A Question Answering Dataset about Sports)
Sport is one of the most popular and revenue-generating forms of entertainment.
1 paper · 0 benchmarks
QDSD (Quantum Dots Stability Diagrams)
This Quantum Dots Stability Diagrams (QDSD) Dataset aggregates experimental stability diagrams of quantum dots from different research groups.
1 paper · 0 benchmarks
A mapping of Quasimodo to the relations of ConceptNet.
1 paper · 0 benchmarks
The R1-Onevision dataset is a meticulously crafted resource designed to empower models with advanced multimodal reasoning capabilities.
1 paper · 0 benchmarks
We conducted a large crowdsourcing study of click patterns in an interactive segmentation scenario and collected 475K real-user clicks.
1 paper · 0 benchmarks
REBUS (A Robust Evaluation Benchmark of Understanding Symbols)
Recent advances in large language models have led to the development of multimodal LLMs (MLLMs), which take both image data and text as an input.
1 paper · 1 benchmark
REFCAT (Internet Archive Scholar Reference Dataset)
Internet Archive Scholar Reference Dataset.
1 paper · 0 benchmarks
Reader Emotion News 20k Dataset
1 paper · 0 benchmarks
RES-Q (RES-Q: Evaluating Code-Editing Large Language Model Systems at the Repository Scale)
RES-Q is a natural language instruction-based benchmark for evaluating Repository Editing Systems, which consists of 100 handcrafted repository editing tasks derived from real GitHub commits.
1 paper · 1 benchmark
Asthma is a common, usually long-term respiratory disease with negative impact on society and the economy worldwide.
1 paper · 0 benchmarks
RFSD (Russian Financial Statements Database)
The Russian Financial Statements Database (RFSD) The Russian Financial Statements Database (RFSD) is an open, harmonized collection of annual unconsolidated financial statements of the universe of Russian firms.
1 paper · 0 benchmarks
RGZ EMU: Semantic Taxonomy (Radio Galaxy Zoo EMU: Towards a Semantic Radio Galaxy Morphology Taxonomy)
The data used in - "Radio Galaxy Zoo EMU: Towards a Semantic Radio Galaxy Morphology Taxonomy" (Bowles et al.
1 paper · 0 benchmarks
The datasets of "Reinforcement Learning-enhanced Shared-account Cross-domain Sequential Recommendation" (TKDE 2022)
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
RLM25 (Research-Level Mathematics 2025)
RLM25 is an evaluation benchmark containing 619 paired examples of research-level natural language mathematical statements and their corresponding Lean formalizations.
1 paper · 0 benchmarks
ROAST (Review level Opinion Aspect Sentiment Target Joint Detection for ABSA)
This repository has a review-level multidomain multilingual dataset for Aspect-based Sentiment Analysis(ABSA) for the paper ROAST: Review-level Opinion Aspect Sentiment Target Joint Detection.
1 paper · 0 benchmarks
RPCD (Reddit Photo Critique Dataset)
The Reddit Photo Critique Dataset (RPCD) contains tuples of image and photo critiques.
1 paper · 0 benchmarks
RPEval (Role-Playing Evaluation Dataset)
Role-Playing Eval (RPEval) is a benchmark dataset designed to evaluate large language models' role-playing abilities across emotional understanding, decision-making, moral alignment, and in-character consistency.
1 paper · 0 benchmarks
The following files contains the simulation inputs and outputs for conducting the multi-objetive optimization of thermal comfort and dyalight with the Response Surface Methodology.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
RTB (Robot Tracking Benchmark)
The Robot Tracking Benchmark (RTB) is a synthetic dataset that facilitates the quantitative evaluation of 3D tracking algorithms for multi-body objects.
1 paper · 1 benchmark
Multilingual explainable fact-checking dataset on Russia-Ukraine Conflict 2022
1 paper · 0 benchmarks
RUSS (Rapid Universal Support Service) is a dataset that consists of a collection of 741 real-world step-by-step natural language instructions (raw and annotated) from the open web, and for each: its corresponding webpage DOM, ground-truth…
1 paper · 0 benchmarks
RVL-CDIP_MP (RVL-CDIP multi-page)
RVL-CDIPMP is our first contribution to retrieve the original documents of the IIT-CDIP test collection which were used to create RVL-CDIP.
1 paper · 0 benchmarks
RVL-CDIP_N_MP (RVL-CDIP-N multi-page)
RVL-CDIPMP-N can serve its original goal as a covariate shift test set, now for multi-page document classification.
1 paper · 0 benchmarks
RWD-10K (Rogue Wave Dataset-10K)
Rogue Wave Dataset-10K dataset consists of 10191 rogue wave images.
1 paper · 0 benchmarks
The data set contains multimodal sensor data generated by a tracked mobile robot in an outdoor and an indoor environemnt.
1 paper · 0 benchmarks
To effectively measure the alignment between automatic evaluation metrics and radiologists' assessments in medical text generation tasks, we have established a comprehensive benchmark, RaTE-Eval, that encompasses three tasks, each with its…
1 paper · 0 benchmarks
RaTE-NER dataset is a large-scale, radiological named entity recognition (NER) dataset, including 13,235 manually annotated sentences from 1,816 reports within the MIMIC-IV database, that spans 9 imaging modalities and 23 anatomical…
1 paper · 0 benchmarks
RadCases Dataset This HuggingFace (HF) dataset contains the raw case labels for input patient "one-liner" case summaries according to the ACR Appropriateness Criteria.
1 paper · 0 benchmarks
The ROAD dataset is made up of observations from the Low Frequency Array (LOFAR) telescope.
1 paper · 0 benchmarks
A total of 227 cross sectional images (20 x 54 mm with a resolution of 289 x 648 pixels) of hind-leg xenograft tumors from 29 mice were obtained with 1mm step-wise movement of the array mounted on a manual positioning device.
1 paper · 0 benchmarks
RadioTalk is a corpus of speech recognition transcripts sampled from talk radio broadcasts in the United States between October of 2018 and March of 2019.
1 paper · 0 benchmarks
The dataset contains generated random signals for autoencoding purposes.
1 paper · 0 benchmarks
The NMR-POISE paper can be found at: Anal.
1 paper · 0 benchmarks
Raw-Microscopy: 940 raw bright-field microscopy images of human blood smear slides for leukocyte classification (microscopy/images/rawscale100) with corresponding labels (microscopy/labels).
1 paper · 0 benchmarks
Synthetic humans generated by the RePoGen method.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.