Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 56 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 2641–2688 of 3,998
FreeMan is the first large-scale multi-view human motion dataset under real scenarios.
1 paper · 0 benchmarks
Fruits Dataset for Classification About Dataset (strawberries, peaches, pomegranates) Photo requirements: 1-White background 2-.jpg 3- Image size 300300 The number of photos required is 250 photos of each fruit when it is fresh and 250…
1 paper · 0 benchmarks
- The dataset contains full-spectral autofluorescence lifetime microscopic images (FS-FLIM) acquired on unstained ex-vivo human lung tissue, where 100 4D hypercubes of 256x256 (spatial resolution) x 32 (time bins) x 512 (spectral channels…
1 paper · 0 benchmarks
FullTextPeerRead is a dataset created by Jeong et al.
1 paper · 0 benchmarks
Demonstration data for 4 FurnitureBench tasks collected with a SpaceMouse using a DiffIK Controller.
1 paper · 0 benchmarks
This dataset was created to test whether it's possible to build a general-purpose detector that can tell real images apart from fake ones generated by convolutional neural networks (CNNs), no matter which model or dataset was used to…
1 paper · 0 benchmarks
GD-NLI (Generated Debiased NLI Datasets)
This is a set of debiased Natural Language Inference (NLI) datasets produced by the paper Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets.
1 paper · 0 benchmarks
The GDIT Aerial Airport dataset consists of aerial images containing instances of parked airplanes.
1 paper · 0 benchmarks
Dataset Card for Dataset Name This dataset is a filtered version of BookCorpus containing only gender-neutral words.
1 paper · 0 benchmarks
GENTER (GEnder Name TEmplates with pRonouns)
This dataset consists of template sentences associating first names ([NAME]) with third-person singular pronouns ([PRONOUN]), e.g., [NAME] asked , not sounding as if [PRONOUN] cared about the answer .
1 paper · 0 benchmarks
This dataset contains short sentences linking a first name, represented by the template mask [NAME], to stereotypical associations.
1 paper · 0 benchmarks
This repository is an extension of GEval.
1 paper · 0 benchmarks
GF-PA66 3D XCT (Glass fiber-reinforced polyamide 66 (GF-PA66) 3D X-ray Computed Tomography (XCT)))
Stack of 2D gray images of glass fiber-reinforced polyamide 66 (GF-PA66) 3D X-ray Computed Tomography (XCT) specimen.
1 paper · 1 benchmark
Stack of 2D gray images of glass fiber-reinforced polyamide 66 (GF-PA66) 3D X-ray Computed Tomography (XCT) specimen.
1 paper · 0 benchmarks
The released GIF Reply dataset contains 1,562,701 real text-GIF conversation turns on Twitter.
1 paper · 1 benchmark
A random sample of 200 machine learning publications, systematically analyzed by a team of labelers, who asked up to 15 questions about how the publication discusses its training data.
1 paper · 0 benchmarks
data/images: data/images/Base : 132 screenshots of game1 & game2 with UI display issues from 466 test reports.
1 paper · 0 benchmarks
A dataset for medical consultation dialogues.
1 paper · 0 benchmarks
GMSC (Give Me Some Credit)
Data for a Kaggle competition Banks play a crucial role in market economies.
1 paper · 0 benchmarks
GO21 is a biomedical knowledge graph that models genes, proteins, drugs, and the hierarchy of the biological processes they participate in.
1 paper · 1 benchmark
GPR-bench (General‑Purpose Reproducibility Benchmark)
GPR‑bench is an open‑source, multilingual benchmark for regression testing and reproducibility tracking in generative‑AI systems.
1 paper · 0 benchmarks
GPTKB is a large general-domain knowledge base (KB) constructed entirely from a large language model (LLM).
1 paper · 0 benchmarks
This is the static test data from the study "Global Geolocated Realtime Data of Interfleet Urban Transit Bus Iding" collected by GRD-TRT-BUF-4I.
1 paper · 0 benchmarks
A scholarly named entity recognition dataset with focus on machine learning models and datasets.
1 paper · 0 benchmarks
GTA-UAV dataset provides a large continuous area dataset (covering 81.3km2) for UAV visual geo-localization, expanding the previously aligned drone-satellite pairs to arbitrary drone-satellite pairs to better align with real-world…
1 paper · 0 benchmarks
GUISS dataset (Meshes, textures, Blend files, stereo datasets, depth maps, depth estimations))
We provide all the expected data inputs to GUISS such as meshes, texture images, and blend files.
1 paper · 0 benchmarks
Details about the creation of the dataset can be seen in https://arxiv.org/abs/2110.06139.
1 paper · 0 benchmarks
GameQA is a large-scale, diverse, and challenging multimodal reasoning dataset designed to enhance the general reasoning capabilities of Vision Language Models (VLMs).
1 paper · 0 benchmarks
GameWikiSum is a domain-specific (video game) dataset for multi-document summarization, which is one hundred times larger than commonly used datasets, and in another domain than news.
1 paper · 0 benchmarks
We construct Gaze-CIFAR-10, a gaze-augmented image dataset based on the standard CIFAR-10 benchmark, enhanced with human eye-tracking annotations collected using the HTC VIVE Pro Eye headset.
1 paper · 1 benchmark
GeBiD (Geometric shapes Bimodal Dataset)
We provide a custom synthetic bimodal dataset, called GeBiD, designed specifically for the comparison of the joint- and cross-generative capabilities of Multimodal Variational Autoencoders.
1 paper · 0 benchmarks
GelSight Young's Modulus Dataset ============== by Michael Burgess Dataset of tactile images collected over grasping common objects labelled with the objects' Young's Moduli.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
GenAIPABench is a specialized dataset designed to evaluate Generative AI-based Privacy Assistants (GenAIPAs).
1 paper · 0 benchmarks
GenPlot (GenPlot: 500k pre-generated plots)
This dataset contains the pre-generated dataset referenced in the GenPlot Paper.
1 paper · 0 benchmarks
https://osf.io/btjqw/?viewonly=f31eda86e7b04ac886734a26cd2ce43d
1 paper · 0 benchmarks
Genomics Adversarial Attack Sample dataset
1 paper · 0 benchmarks
GenoTEX (An LLM Agent Benchmark for Automated Gene Expression Data Analysis)
GenoTEX (Genomics Data Automatic Exploration Benchmark) is a benchmark dataset for the automated analysis of gene expression data to identify disease-associated genes while considering the influence of other biological factors.
1 paper · 0 benchmarks
GeoJEPAD is a multimodal dataset combining OpenStreetMap (OSM) data (attributes and geometries) with high-resolution aerial imagery from diverse urban areas.
1 paper · 0 benchmarks
GeoQuestions1089 is a crowdsourced geospatial question-answering dataset that targets the Knowledge Graph YAGO2geo.
1 paper · 1 benchmark
Dataset Description This dataset contains rental property listings scraped from Tonaton.com, one of Ghana's leading online classifieds platforms.
1 paper · 0 benchmarks
The Gigaword Entailment dataset is a dataset for entailment prediction between an article and its headline.
1 paper · 0 benchmarks
GitBugs (GitBugs: Bug Reports for Duplicate Detection, Retrieval Augmented Generation, Triage, and More)
GitBugs is a comprehensive and up-to-date dataset comprising over 150,000 bug reports from nine actively maintained open-source projects, including Firefox, Cassandra, and VS Code.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Description This Dataset contains review information on Google map (ratings, text, images, etc.), business metadata (address, geographical info, descriptions, category information, price, open hours, and MISC info), and links (relative…
1 paper · 0 benchmarks
This dataset was curated for Search Engine Optimization (SEO) analysis tasks, including categorization and spam detection.
1 paper · 0 benchmarks
Dataset introduced by Xifeng Yan et al.
1 paper · 0 benchmarks
Dataset introduced by Xifeng Yan et al.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.