Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 188 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 8977–9024 of 12,172
Images collected on an LED array microscope (also known as a Fourier ptychographic microscope) on 172 fields-of-view of frog blood smears.
1 paper · 0 benchmarks
The LEMMA dataset aims to explore the essence of complex human activities in a goal-directed, multi-agent, multi-task setting with ground-truth labels of compositional atomic-actions and their associated tasks.
1 paper · 0 benchmarks
LEMONADE is a large, expert-annotated dataset for event extraction from news articles in 20 languages: English, Spanish, Arabic, French, Italian, Russian, German, Turkish, Burmese, Indonesian, Ukrainian, Korean, Portuguese, Dutch, Somali,…
1 paper · 0 benchmarks
This data set comprises 22 fundus images with their corresponding manual annotations for the blood vessels, separated as arteries and veins.
1 paper · 2 benchmarks
LGI-PPGI is a dataset for heart Rate estimation from face videos in the wild.
1 paper · 0 benchmarks
Generated for further pre-training pre-trained models like BERT, RoBERTa, ALBERT, DeBERTa, etc..
1 paper · 0 benchmarks
LIB-HSI (RGB and Hyperspectral images of Building Facades)
The LIB-HSI dataset contains hyperspectral reflectance images and their corresponding RGB images of building façades in a light industrial environment.
1 paper · 0 benchmarks
A multimodal LIBRAS-UFOP Brazilian sign language dataset of minimal pairs using a microsoft Kinect senso.
1 paper · 1 benchmark
The National Institute of Informatics provides LIFULL HOME'S Dataset to researchers, which was offered by LIFULL Co., Ltd.
1 paper · 0 benchmarks
LIGHT-Quests is an extension of LIGHT, a large-scale crowd-sourced fantasy text-game, to generate a dataset of quests.
1 paper · 0 benchmarks
LIRCAD (Inria Liver vessels subbranch anotomical nomenclature labels - "LIRCAD")
The structure for the dataset is as follows : 3DLiverVasculatureProject/ ├── CT/ │ Contains the CT scans ├── Labels/ │ Contains the dual labels for the vessel tree annotations 0 background, 1 Portal vein, 2 hepatic vein ├──…
1 paper · 0 benchmarks
The LIRIS human activities dataset contains (gray/rgb/depth) videos showing people performing various activities taken from daily life (discussing, telphone calls, giving an item etc.).
1 paper · 0 benchmarks
LISA Gaze is a dataset for driver gaze estimation comprising of 11 long drives, driven by 10 subjects in two different cars.
1 paper · 0 benchmarks
LIV360SV (Liverpool 360 degree Street View)
The dataset contains 26,645, 360 degree, street-level images collected via cycling with a GoPro Fusion camera, recorded Jan 14th -- 18th 2020.
1 paper · 0 benchmarks
This dataset comprises high-quality, targeted spear-phishing emails created using a proprietary system that harnesses the power of LLMs and knowledge graphs.
1 paper · 0 benchmarks
LLM Health Benchmarks Dataset The Health Benchmarks Dataset is a specialized resource for evaluating large language models (LLMs) in different medical specialties.
1 paper · 0 benchmarks
Dataset is a CSV file, that contains evaluation scores given by a panel of LLMs to responses produced by other LLMs .
1 paper · 0 benchmarks
To evaluate our proposed strategy of asynchronous communication for LLMs, we run games of Mafia with human players, incorporating an LLM-based agent as an additional player, within an asynchronous chat environment.
1 paper · 0 benchmarks
Three tasks were addressed in the LLMs4OL paradigm.
1 paper · 0 benchmarks
LLNeRF Dataset is a real-world dataset as a benchmark for model learning and evaluation.
1 paper · 0 benchmarks
LLaVA-Rad MIMIC-CXR features more accurate section extractions from MIMIC-CXR free-text radiology reports.
1 paper · 0 benchmarks
Are Large Pre-Trained Language Models Leaking Your Personal Information?
1 paper · 0 benchmarks
A diverse set of 21 relations, each covering a different set of subject-entities and a complete list of ground truth object-entities per subject-relation-pair.
1 paper · 1 benchmark
The LM-O (Linemod-Occluded) dataset, introduced by Brachmann et al.
1 paper · 0 benchmarks
LMC (Language Model Council)
The Language Model Council (LMC) is a novel benchmarking framework proposed to address the challenge of ranking Large Language Models (LLMs) on highly subjective tasks¹.
1 paper · 0 benchmarks
LMOT (Low-light Multi-object Tracking Dataset)
The Low-light Multi-object Tracking Dataset (LMOT) is a large-scale dataset that focuses on multi-object tracking in dark scenes.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
While stroking a rigid tool over an object surface, vibrations induced on the tool, which represent the interaction between the tool and the surface texture, can be measured by means of an accelerometer.
1 paper · 0 benchmarks
LPBA40 (LONI Probabilistic Brain Atlas)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
LPSC (Planetary Science Data Set)
This data set contains annotated text versions of 1635 two-page abstracts published at the Lunar and Planetary Science Conference from 1998 to 2020 of relevance to four Mars missions.
1 paper · 2 benchmarks
This is a dataset is composed of full-document images, groundtruth, and tools to perform an evaluation of binarization algorithms.
1 paper · 0 benchmarks
LSA-T (Lengua de Señas Argentina - Traducción)
LSA-T is the first continuous Argentinian Sign Language (LSA) dataset.
1 paper · 1 benchmark
LSDBench (Long-video Sampling Dilemma Benchmark)
A benchmark that focuses on the sampling dilemma in long-video tasks.
1 paper · 0 benchmarks
LSEC (Live Stream E-Commerce)
The LSEC (Live Stream E-Commerce) dataset has two subsets: LSEC-Small and LSEC-Large.
1 paper · 0 benchmarks
Sign Language Datasets for French Belgian Sign Language This dataset is built upon the work of Belgian linguists from the University of Namur.
1 paper · 0 benchmarks
LSICC (Large Scale Informal Chinese Corpus)
Large Scale Informal Chinese Corpus (LSICC) is a large-scale corpus of informal Chinese.
1 paper · 0 benchmarks
LSLF (Large-scale Labeled Face)
Consists of a large number of unconstrained multi-view and partially occluded faces.
1 paper · 0 benchmarks
The Large Scale Movie Description Challenge (LSMDC) - Context is an augmented version of the original LSMDC dataset with movie scripts as contextual text.
1 paper · 0 benchmarks
LTFT (Long-Term Face Tracking)
Dataset originally conceived for multi-face tracking/detection for highly crowded scenarios.
1 paper · 0 benchmarks
The LTI LangID Corpus is a dataset used for language identification (LangID) tasks.
1 paper · 0 benchmarks
LVVO (Lecture Video Visual Objects)
The Lecture Video Visual Objects (LVVO) dataset is a benchmark designed for object detection in lecture video frames.
1 paper · 0 benchmarks
This dataset consists of more than 16,000 retinal OCT B-scans from 441 cases (Normal: 120, Drusen: 160, CNV: 161) and is acquired at Noor Eye Hospital, Tehran, Iran.
1 paper · 0 benchmarks
Citations are an important part of scientific papers, and the proper handling of them is indispensable for the science of science.
1 paper · 0 benchmarks
The whole UCF-Crime dataset consists of real-world 240 × 320 RGB videos with 13 realistic anomaly types such as explosion, road accident, burglary, etc., and normal examples.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Introduction These audio files accompany the preprint by Accolti (2025), which presents a preliminary study on the effect of the acoustical conditions of three different rooms on the perception of virtual stages for music.
1 paper · 0 benchmarks
This data set contains weekly scans of cauliflower and broccoli covering a ten week growth cycle from transplant to harvest.
1 paper · 0 benchmarks
Landscape Dataset consists of landscape images collected from Flickr.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.