Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 55 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 2593–2640 of 3,998
The EyeInfo Dataset is an open-source eye-tracking dataset created by Fabricio Batista Narcizo, a research scientist at the IT University of Copenhagen (ITU) and GN Audio A/S (Jabra), Denmark.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
FABSA (An aspect-based sentiment analysis dataset of Customer Feedback reviews)
FABSA, An aspect-based sentiment analysis dataset in the Customer Feedback space (Trustpilot, Google Play and Apple Store reviews).
1 paper · 2 benchmarks
FB15K237-Refined is a refined version of FB15k237 by KGRefiner.
1 paper · 0 benchmarks
FCoT (Foreground Chain-of-Thought)
FCoT (Chain‑of‑Thought Segmentation) is replicate the step-by-step reasoning process a human annotator follows when using SAM2 to generate masks.
1 paper · 0 benchmarks
A large-scale isolated Indian sign language dataset.
1 paper · 1 benchmark
The data set contains point cloud data captured in an indoor environment with precise localization and ground truth mapping information.
1 paper · 0 benchmarks
Tables of the blendshapes from a group of the images of the FER2013 dataset, generated using MediaPipe library, based on the ARKit face blendshapes.
1 paper · 0 benchmarks
FETA Car-Manuals (FETA Car-Manuals dataset, image-text retrieval for foundation models' expert data performance.)
FETA benchmark focuses on text-to-image and image-to-text retrieval in public car manuals and sales catalogue brochures.
1 paper · 2 benchmarks
FETA benchmark focuses on text-to-image and image-to-text retrieval in public car manuals and sales catalogue brochures.
1 paper · 0 benchmarks
The development of the remote sensing fine-grained ship classification field necessitates large-scale realistic fine-grained ship datasets.
1 paper · 0 benchmarks
FGraDA (Fine-Grained Domain Adaptation Dataset)
Previous research for adapting a general neural machine translation (NMT) model into a specific domain usually neglects the diversity in translation within the same domain, which is a core problem for domain adaptation in real- world…
1 paper · 0 benchmarks
FICLE (Factual Inconsistency CLassification with Explanation)
The FICLE dataset is a derivative of the FEVER dataset, which is a collection of 185,445 claims generated by modifying sentences obtained from Wikipedia.
1 paper · 0 benchmarks
Optical images of printed circuit boards as well as detailed annotations of any text, logos, and surface-mount devices (SMDs).
1 paper · 0 benchmarks
FIND (Fused Image dataset for convolutional neural Network-based crack Detection)
The “Fused Image dataset for convolutional neural Network-based crack Detection” (FIND) is a large-scale image dataset with pixel-level ground truth crack data for deep learning-based crack segmentation analysis.
1 paper · 0 benchmarks
FLIP includes several benchmark datasets that contain a variety of protein sequences, each with a real-valued label indicating its "fitness" (how well the protein performs some particular function).
1 paper · 0 benchmarks
FMC-MWO2KG (The MWO2KG Failure Mode Classification Dataset)
The Failure Mode Classification dataset released in the paper "MWO2KG and Echidna: Constructing and exploring knowledge graphs from maintenance data" by Stewart et al.
1 paper · 1 benchmark
FP4S (Floor plan image segmentation via scribble-based semi-weakly-supervised learning)
We introduce a new style- and category-agnostic floor plan image parsing benchmark developed in collaboration with professional architectural designers.
1 paper · 1 benchmark
FQ-160 (Forbidden Question Dataset (160))
The forbidden question dataset they build (based on two previous works) contains 160 questions from 160 violated categories.
1 paper · 0 benchmarks
FSC-P2 (Fearless Steps Challenge Phase2)
The Fearless Steps Initiative by UTDallas-CRSS led to the digitization, recovery, and diarization of 19,000 hours of original analog audio data, as well as the development of algorithms to extract meaningful information from this…
1 paper · 0 benchmarks
FSOCO is a collaborative dataset for vision-based cone detection systems in Formula Student Driverless competitions.
1 paper · 0 benchmarks
FTR-18 is a multilingual rumour dataset on football transfer news.
1 paper · 0 benchmarks
A set of 248 search queries annotated with the correct diagnosis.
1 paper · 0 benchmarks
Facial Skeletal angles (Facial Skeletal Angles (Glabella and Maxilla Angle and Length and Width of Piriformis))
Facial Skeletal Angles (Glabella and Maxilla Angle and Length and Width of Piriformis)
1 paper · 0 benchmarks
The FairTranslate Dataset includes 2,418 sentence pairs, each centered around an occupation, designed to assess gender expression and translation in English-French contexts.
1 paper · 0 benchmarks
Fallout New Vegas Dialog is a multilingual sentiment annotated dialog dataset from Fallout New Vegas.
1 paper · 0 benchmarks
The "Famous Keyword Twitter Replies Dataset" is a comprehensive collection of Twitter data that focuses on popular keywords and their associated replies.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
FedNLP (FOMC Docs and Speeches)
We collect the various forms of Federal Reserve communications.
1 paper · 0 benchmarks
DOI: https://doi.org/10.7910/DVN/O4CRXK The most comprehensive standardised data on Malaysian federal and state elections from 1955 to the present.
1 paper · 0 benchmarks
Introduction The FewGLUE64labeled dataset is a new version of FewGLUE dataset.
1 paper · 0 benchmarks
FiVE (A Fine-grained Video Editing Benchmark)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This dataset is a "part II" extension of the "Engineered cardiac microbundle time-lapse microscopy image dataset" and contains 808 experimental time-lapse image sequences of beating hiPSC-based cardiac microbundles using FibroTUG platforms…
1 paper · 0 benchmarks
FinDKG: The Global Financial Dynamic Knowledge Graph Dataset FinDKG is an open-source dataset focused on creating a temporally-resolved Financial Dynamic Knowledge Graph.
1 paper · 0 benchmarks
Financial Language Understanding Evaluation is an open-source comprehensive suite of benchmarks for the financial domain.
1 paper · 0 benchmarks
FineCops-Ref is a dataset for Compositional Referring Expression Comprehension (REC) that rigorously evaluates Vision-Language Models (VLMs) on compositional reasoning and their ability to identify inconsistencies between images and text.
1 paper · 0 benchmarks
Synthetic training set: This set is constructed in the following two steps and will be used for estimation/training purposes.
1 paper · 0 benchmarks
The researchers collected a dataset of 3,500 images of Tilapia fish in a small bowl containing three fish per image.
1 paper · 0 benchmarks
The researchers collected 3,500 images of Tilapia fish, with each image containing three fish in a small bowl.
1 paper · 0 benchmarks
FloCo (Flow chart Image to Code)
the FloCo dataset that contains 11,884 flowchart images and their corresponding Python codes.
1 paper · 1 benchmark
About Dataset The file contains 24K unique figure obtained from various Google resources Meticulously curated figure ensuring diversity and representativeness Provides a solid foundation for developing robust and precise figure allocation…
1 paper · 0 benchmarks
The Food Recall Incidents dataset consists of 7,546 short texts (from 5 to 360 characters each), which are the titles of food recall announcements (therefore referred to as title), crawled from 24 public food safety authority websites by…
1 paper · 0 benchmarks
Food.com Recipes and Interactions consists of 270K recipes and 1.4M user-recipe interactions (reviews) scraped from Food.com, covering a period of 18 years (January 2000 to December 2018).
1 paper · 0 benchmarks
Repository for the question sets and resolution sets described produced by ForecastBench, a forecasting benchmark for LLMs.
1 paper · 0 benchmarks
This dataset contains news headlines relevant to key forex pairs: AUDUSD, EURCHF, EURUSD, GBPUSD, and USDJPY.
1 paper · 0 benchmarks
The 'Me 163' was a Second World War fighter airplane and a result of the German air force secret developments.
1 paper · 0 benchmarks
The Fraunhofer Portugal AICOS EDoF Dataset was produced within the TAMI project and is composed of images of microscopic fields of view (FOV) of Liquid-based Cervical Cytology (LBC) samples.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.