Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 244 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 11665–11712 of 12,172
Dataset Description: Summarized Wiki Articles with TTL Knowledge Graphs Overview This dataset comprises 500 summarized Wikipedia articles, each accompanied by a corresponding TTL knowledge graph.
0 papers · 0 benchmarks
KID-F (K-pop Idol Dataset - Female)
Description K-pop Idol Dataset - Female (KID-F) is the first dataset of K-pop idol high quality face images.
0 papers · 0 benchmarks
The KIT Whole-Body Human Motion Database is a large-scale dataset of whole-body human motion with methods and tools, which allows a unifying representation of captured human motion, and efficient search in the database, as well as the…
0 papers · 0 benchmarks
KITTI-360-SR (KITTI-360 modification for Scene Recognition task)
Scene Recognition is a problem, where a set of visible objects must be correctly associated with objects marked on a semantic map - this problem is also sometimes called a Data Association.
0 papers · 0 benchmarks
The KITTI-Motion dataset contains pixel-wise semantic class labels and moving object annotations for 255 images taken from the KITTI Raw dataset.
0 papers · 0 benchmarks
KTI Multiview Football I is a dataset of football players with annotated joints that can be used for multi-view reconstruction.
0 papers · 0 benchmarks
KinFaceW consists of two kinship datasets: KinFaceW-I and KinFaceW-II.
0 papers · 0 benchmarks
To study kinship verification from gait, we collected the dataset KinGaitWild consisting of several videos from youtube.
0 papers · 0 benchmarks
The Dataset consists of the multimodal facial images of 52 people (14 females, 38 males) obtained by Kinect.
0 papers · 0 benchmarks
L-SVD (Large-Scale Selfie Video Dataset (L-SVD): A Benchmark for Emotion Recognition)
Welcome to L-SVD L-SVD is an extensive and rigorously curated video dataset aimed at transforming the field of emotion recognition.
0 papers · 0 benchmarks
The L1000 dataset consists of ~1,400,000 gene-expression profiles on the responses of ~50 human cell lines to one of ~20,000 compounds across a range of concentrations.
0 papers · 0 benchmarks
LASIESTA (Labeled and Annotated Sequences for Integral Evaluation of SegmenTation Algorithms) is a segmentation and detection dataset composed by many real indoor and outdoor sequences organized into categories, each of one covering a…
0 papers · 0 benchmarks
This dataset is suitable for sentiment analysis, which consists of Danish data from the Leipzig Collection.
0 papers · 0 benchmarks
This is the synthetic dataset used for training a model which alerts users for potential leakages of personal information.
0 papers · 0 benchmarks
This is a dataset for vehicle detection.
0 papers · 0 benchmarks
LLSD (Low-illumination steel dataset)
A dataset of low-illumination steel for industrial applications
0 papers · 0 benchmarks
LMCQA (Legal Multiple Choice Question Answering)
This dataset contains a set of multiple-choice questions related to various legal topics.
0 papers · 0 benchmarks
The Landsat collection contains 400x400 RGB pictures captured by the Landsat 8 satellite.
0 papers · 0 benchmarks
LRWC (Lexical Relations from the Wisdom of the Crowd)
This dataset contains the opinions of Russian native speakers about the relationship between a generic term (hypernym) and a specific instance of it (hyponym).
0 papers · 0 benchmarks
LSARS (Large Scale Abstractive multi-Review Summarization)
In an active e-commerce environment, customers process a large number of reviews when deciding on whether to buy a product or not.
0 papers · 0 benchmarks
LSFM (Large Scale Facial Model (LSFM))
The Large Scale Facial Model (LSFM) is a 3D statistical model of facial shape built from nearly 10,000 individuals.
0 papers · 0 benchmarks
LTIR (Linköping Thermal InfraRed)
The LTIR dataset is a thermal infrared dataset for evaluation of Short-Term Single-Object (STSO) tracking.
0 papers · 0 benchmarks
A landslide dataset for LDBF predictions.
0 papers · 0 benchmarks
LeQua2024 (Learning to Quantify Dataset 2024)
This is the dataset used in the 2nd data challenge on Learning to Quantify.
0 papers · 0 benchmarks
Court decisions from 2017 and 2018 were selected for the dataset, published online by the Federal Ministry of Justice and Consumer Protection.
0 papers · 0 benchmarks
Lemon dataset has been prepared to investigate the possibilities to tackle the issue of fruit quality control.
0 papers · 0 benchmarks
LfED-6D (Learning from Experience and Demonstration for 6-DOF Grasping Dataset)
The LfED-6D dataset is a collection of 6D grasp annotations acquired through experience (with a robot platform) or by human demonstration.
0 papers · 0 benchmarks
LiMiT (Literal Motion in Text Dataset)
The limit dataset of ~24K sentences that describe literal motion (~14K sentences), and sentences not describing motion or other type of motion (e.g.
0 papers · 0 benchmarks
LiSu (LiSu: A Dataset and Method for LiDAR Surface Normal Estimation)
We present LiSu, a novel synthetic LiDAR dataset targeted for research on surface normal estimation.
0 papers · 0 benchmarks
The LinkSO dataset is a resource for learning to retrieve similar question-answer pairs on Stack Overflow.
0 papers · 0 benchmarks
We introduce low-light image enhancement benchmark dataset “Low-light Images of Streets (LoLI-Street),” which contains three subsets: train, validation, and test.
0 papers · 0 benchmarks
The Lusitano dataset was collected over a 3-month period, spanning from January to March, from Paulo de Oliveira, S.A., a prominent textile company, based in Covilhã, Portugal, renowned for its innovative contributions to the textile…
0 papers · 0 benchmarks
MAEC (Multimodal Aligned Earnings Conference Call Dataset)
MAEC is a new, large-scale multi-modal, text-audio paired, earnings-call dataset named MAEC, based on S&P 1500 companies.
0 papers · 0 benchmarks
MASATI (MAritime SATellite Imagery dataset)
The MASATI dataset contains color images in dynamic marine environments, and it can be used to evaluate ship detection methods.
0 papers · 0 benchmarks
A dataset for multi-context visual grounding.
0 papers · 0 benchmarks
MCCSD (Mandarin Chinese Cued Speech Dataset)
This MCCS dataset is the first large-scale Mandarin Chinese Cued Speech dataset.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.