Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 59 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 2785–2832 of 3,998
JustLogic is a natural language deductive reasoning dataset.
1 paper · 0 benchmarks
This is the dataset for testing the robustness of various VO/VIO methods, acquired on reak UAV.
1 paper · 0 benchmarks
KCIF (Knowledge Conditioned Instruction Following (KCIF))
KCIF is a benchmark for evaluating the instruction-following capabilities of Large Language Models (LLM).
1 paper · 0 benchmarks
KGRED (Knowledge-graph-enhanced relation extraction datasets--)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Extension of the official KITTI'15 dataset.
1 paper · 0 benchmarks
KITTI-6DoF is a dataset that contains annotations for the 6DoF estimation task for 5 object categories on 7,481 frames.
1 paper · 0 benchmarks
Please find more details of this dataset at https://alex-xun-xu.github.io/ProjectPage/CVPR18/index.html 3D motion segmentation has been the key problem in computer vision research due to the application in structure from motion and…
1 paper · 1 benchmark
At KayifamilyTv, we introduce a powerful and scalable video-mining solution designed to enhance captioning capabilities by transferring supervision from image datasets to video and audio content.
1 paper · 0 benchmarks
Kinetics-GEB+ (Generic Event Boundary Captioning, Grounding and Retrieval) is a dataset that consists of over 170k boundaries associated with captions describing status changes in the generic events in 12K videos.
1 paper · 3 benchmarks
The Kite database is a multi-modal dataset for the control of unmanned aerial vehicles (UAVs).
1 paper · 0 benchmarks
We introduce KnowledJe, an English-language knowledge graph of antisemitic history and language from the 20th century to the present.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Full Koo platform Dataset.
1 paper · 0 benchmarks
a parallel corpus of Sorani (ckb or Central Kurdish) and Kurmanji (kmr or Northern Kurdish) dialects of Kurdish along with English (eng).
1 paper · 0 benchmarks
Cleaned and preprocessed version of the Kyokushin Karate Motion Dataset by Szczkesna et al.
1 paper · 0 benchmarks
The Sentinel-2 satellite carries 12 CMOS detectors for the VNIR bands, with adjacent detectors having overlapping fields of view that result in overlapping regions in level-1 B (L1B) images.
1 paper · 0 benchmarks
Raw negotiation transcripts generated for the paper "Evaluating Language Model Agency through Negotiations".
1 paper · 0 benchmarks
For our experiments, we collected a dataset of procedural knowledge of the LangChain Python library, unseen by many extant LLMs.
1 paper · 0 benchmarks
The datasets of "Towards Lightweight Cross-domain Sequential Recommendation via External Attention-enhanced Graph Convolution Network" (DASFAA 2023)
1 paper · 0 benchmarks
The dataset was collected from two courses offered on the University of Jordan's E-learning Portal during the second semester of 2020, namely "Computer Skills for Humanities Students" (CSHS) and "Computer Skills for Medical Students"…
1 paper · 0 benchmarks
LEMONADE is a large, expert-annotated dataset for event extraction from news articles in 20 languages: English, Spanish, Arabic, French, Italian, Russian, German, Turkish, Burmese, Indonesian, Ukrainian, Korean, Portuguese, Dutch, Somali,…
1 paper · 0 benchmarks
LIB-HSI (RGB and Hyperspectral images of Building Facades)
The LIB-HSI dataset contains hyperspectral reflectance images and their corresponding RGB images of building façades in a light industrial environment.
1 paper · 0 benchmarks
LIGHT-Quests is an extension of LIGHT, a large-scale crowd-sourced fantasy text-game, to generate a dataset of quests.
1 paper · 0 benchmarks
The LIRIS human activities dataset contains (gray/rgb/depth) videos showing people performing various activities taken from daily life (discussing, telphone calls, giving an item etc.).
1 paper · 0 benchmarks
This dataset comprises high-quality, targeted spear-phishing emails created using a proprietary system that harnesses the power of LLMs and knowledge graphs.
1 paper · 0 benchmarks
LLM Health Benchmarks Dataset The Health Benchmarks Dataset is a specialized resource for evaluating large language models (LLMs) in different medical specialties.
1 paper · 0 benchmarks
Dataset is a CSV file, that contains evaluation scores given by a panel of LLMs to responses produced by other LLMs .
1 paper · 0 benchmarks
To evaluate our proposed strategy of asynchronous communication for LLMs, we run games of Mafia with human players, incorporating an LLM-based agent as an additional player, within an asynchronous chat environment.
1 paper · 0 benchmarks
Three tasks were addressed in the LLMs4OL paradigm.
1 paper · 0 benchmarks
LLNeRF Dataset is a real-world dataset as a benchmark for model learning and evaluation.
1 paper · 0 benchmarks
LLaVA-Rad MIMIC-CXR features more accurate section extractions from MIMIC-CXR free-text radiology reports.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
LPSC (Planetary Science Data Set)
This data set contains annotated text versions of 1635 two-page abstracts published at the Lunar and Planetary Science Conference from 1998 to 2020 of relevance to four Mars missions.
1 paper · 2 benchmarks
LSDBench (Long-video Sampling Dilemma Benchmark)
A benchmark that focuses on the sampling dilemma in long-video tasks.
1 paper · 0 benchmarks
LVVO (Lecture Video Visual Objects)
The Lecture Video Visual Objects (LVVO) dataset is a benchmark designed for object detection in lecture video frames.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Introduction These audio files accompany the preprint by Accolti (2025), which presents a preliminary study on the effect of the acoustical conditions of three different rooms on the perception of virtual stages for music.
1 paper · 0 benchmarks
This data set contains weekly scans of cauliflower and broccoli covering a ten week growth cycle from transplant to harvest.
1 paper · 0 benchmarks
Studying how human drivers react differently when following autonomous vehicles (AV) vs.
1 paper · 0 benchmarks
This dataset contains two types of audio recordings.
1 paper · 0 benchmarks
LeafNet (LeafNet: A large-scale dataset for training image-text models in leaf disease identification)
The PlantVillage dataset, with over 54,000 images spanning 14 plant species and 26 disease types, has been widely used for leaf disease classification.
1 paper · 1 benchmark
Dataset Summary New dataset introduced in Parameter-Efficient Legal Domain Adaptation (Li et al., 2022) from the Legal Advice Reddit community (known as "/r/legaldvice"), sourcing the Reddit posts from the Pushshift Reddit dataset.
1 paper · 0 benchmarks
Recognizing events and their coreferential men- tions in a document is essential for understand- ing semantic meanings of text.
1 paper · 0 benchmarks
This dataset includes sharp-blur pairs of Leishmania image, which is a protozoan parasite microscopy image dataset of Leishmania, obtained from the preserved slides stained with Giemsa.
1 paper · 0 benchmarks
LibriS2S is a Speech to Speech Translation (S2ST) dataset build further upon existing resources.
1 paper · 0 benchmarks
We created this robust and custom light field dataset in order to assist light field researchers in using SOTA machine learning algorithms for a variety of light field tasks such as depth estimation, synthetic aperture imaging, and more.
1 paper · 0 benchmarks
LimeSoDa (Precision Liming Soil Datasets)
Precision Liming Soil Datasets (LimeSoDa) is a collection of 31 datasets from a field- and farm-scale soil mapping context.
1 paper · 0 benchmarks
The Lincolnbeet dataset is an object detection dataset designed to encourage research in the identification of items in environments with high levels of occlusion, and in the development of better approaches to evaluate object detection…
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.