Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 83 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 3937–3984 of 12,172
Egocentric motion capture dataset
6 papers · 1 benchmark
H-DIBCO 2014 is the International Document Image Binarization Competition which is dedicated to handwritten document images organized in conjunction with ICFHR 2014 conference.
6 papers · 0 benchmarks
H-DIBCO 2018 is the international Handwritten Document Image Binarization Contest organized in the context of ICFHR 2018 conference.
6 papers · 0 benchmarks
HRS-Bench (Holistic, Reliable, and Scalable Benchmark)
HRS-Bench is a concrete evaluation benchmark for T2I models that is Holistic, Reliable, and Scalable.
6 papers · 0 benchmarks
HiXray is a High-quality X-ray security inspection image dataset, which contains 102,928 common prohibited items of 8 categories.
6 papers · 0 benchmarks
We introduce HumanEval-XL, a massively multilingual code generation benchmark specifically crafted to address this deficiency.
6 papers · 0 benchmarks
HurricaneEmo is an emotion dataset that contains 15,000 English tweets spanning three hurricanes: Harvey, Irma, and Maria.
6 papers · 0 benchmarks
The IBM-Rank-30k is a dataset for the task of argument quality ranking.
6 papers · 0 benchmarks
iCubWorld datasets are collections of images recording the visual experience of iCub while observing objects in its typical environment, a laboratory or an office.
6 papers · 0 benchmarks
IIIT-AR-13K is created by manually annotating the bounding boxes of graphical or page objects in publicly available annual reports.
6 papers · 0 benchmarks
IIIT-ILST is a dataset and benchmark for scene text recognition for three Indic scripts - Devanagari, Telugu and Malayalam.
6 papers · 0 benchmarks
An IMPlicature and PRESupposition diagnostic dataset (IMPPRES), consisting of >25k semiautomatically generated sentence pairs illustrating well-studied pragmatic inference types.
6 papers · 0 benchmarks
The ISIC 2018 dataset was published by the International Skin Imaging Collaboration (ISIC) as a large-scale dataset of dermoscopy images.
6 papers · 0 benchmarks
The dataset contains 33,126 dermoscopic training images of unique benign and malignant skin lesions from over 2,000 patients.
6 papers · 1 benchmark
The IWSLT 2015 Evaluation Campaign featured three tracks: automatic speech recognition (ASR), spoken language translation (SLT), and machine translation (MT).
6 papers · 0 benchmarks
Dataset of over 6 million GPS-tagged images from Flickr.
6 papers · 1 benchmark
OpenImage-O is built for the ID dataset ImageNet-1k.
6 papers · 1 benchmark
ImageNet-100 is a subset of ImageNet-1k Dataset from ImageNet Large Scale Visual Recognition Challenge 2012.
6 papers · 0 benchmarks
IndustReal (IndustReal Dataset of Egocentric Videos for Procedure Understanding)
IndustReal is an ego-centric, multi-modal dataset where 27 participants are challenged to perform assembly and maintenance procedures on a construction-toy car.
6 papers · 3 benchmarks
A large-scale (105K conversations) media dialog dataset collected from news interview transcripts.
6 papers · 0 benchmarks
JParaCrawl is a parallel corpus for English-Japanese, for which the amount of publicly available parallel corpora is still limited.
6 papers · 0 benchmarks
JerichoWorld is a dataset that enables the creation of learning agents that can build knowledge graph-based world models of interactive narratives.
6 papers · 2 benchmarks
KAMEL (Knowledge Analysis with Multitoken Entities in Language Models)
KAMEL comprises knowledge about 234 relations from Wikidata with a large training, validation, and test dataset.
6 papers · 1 benchmark
Augments the KITTI with more instance pixel-level annotation for 8 categories.
6 papers · 1 benchmark
The task is to predict the chances of a user listening to a song repetitively after the first observable listening event within a time window was triggered.
6 papers · 1 benchmark
KnowledgeNet is a benchmark dataset for the task of automatically populating a knowledge base (Wikidata) with facts expressed in natural language text on the web.
6 papers · 0 benchmarks
Kompetencer (Danish Job Postings Classification Dataset)
Kompetencer (en: competences) is a Danish job posting dataset annotated for nested spans of competences.
6 papers · 0 benchmarks
L3DAS21 is a dataset for 3D audio signal processing.
6 papers · 2 benchmarks
LCQMC (Large-scale Chinese Question Matching Corpus)
LCQMC is a large-scale Chinese question matching corpus.
6 papers · 0 benchmarks
LIS (low-light instance segmentation)
To reveal and systematically investigate the effectiveness of the proposed method in the real world, a real low-light image dataset for instance segmentation is necessary and urgently needed.
6 papers · 0 benchmarks
LIVE (Laboratory for Image & Video Engineering)
Briefly describe the dataset.
6 papers · 2 benchmarks
The unsupervised Labeled Lane MArkerS dataset (LLAMAS) is a dataset for lane detection and segmentation.
6 papers · 1 benchmark
LOGO is a multi-person long-form video dataset with frame-wise annotations on both action procedures and formations based on artistic swimming scenarios.
6 papers · 0 benchmarks
LRD (Low-Light Raw Denoising Dataset)
We collected a new low-light raw denoising (LRD) dataset for training and benchmarking.
6 papers · 0 benchmarks
LSA64 (LSA64: A Dataset for Argentinian Sign Language)
The sign database for the Argentinian Sign Language, created with the goal of producing a dictionary for LSA and training an automatic sign recognizer, includes 3200 videos where 10 non-expert subjects executed 5 repetitions of 64…
6 papers · 1 benchmark
LSMDC-E contains 20,151 training samples, 1,477 validation samples and 2,005 test samples, which is modified from LSMDC 2021.
6 papers · 1 benchmark
LSMI (Large Scale Multi-Illuminant dataet)
Large Scale Multi-Illuminant (LSMI) Dataset for Developing White Balance Algorithm under Mixed Illumination (ICCV 2021) Change Log LSMI Dataset Version : 1.1 1.0 : LSMI dataset released.
6 papers · 0 benchmarks
LSSED, a challenging large-scale english dataset for speech emotion recognition.
6 papers · 1 benchmark
The LastLetterConcat dataset is a collection of word concatenations formed by taking the last letters of individual words and joining them together.
6 papers · 0 benchmarks
LiDAR-MOS (LiDAR-based Moving Object Segmentation)
Tasks.
6 papers · 0 benchmarks
Libri-adhoc40 is a synchronized speech corpus which collects the replayed Librispeech data from loudspeakers by ad-hoc microphone arrays of 40 strongly synchronized distributed nodes in a real office environment.
6 papers · 0 benchmarks
We randomly selected three videos from the Internet, that are longer than 1.5K frames and have their main objects continuously appearing.
6 papers · 1 benchmark
Lyra is a dataset for code generation that consists on Python code with embedded SQL.
6 papers · 0 benchmarks
Lytro Illum is a new light field dataset using a Lytro Illum camera.
6 papers · 0 benchmarks
MARIDA (Marine Debris Archive)
MARIDA (Marine Debris Archive) is the first dataset based on the multispectral Sentinel-2 (S2) satellite data, which distinguishes Marine Debris from various marine features that co-exist, including Sargassum macroalgae, Ships, Natural…
6 papers · 1 benchmark
MCXFACE (Multi-Channel Heterogeneous Face Recognition dataset)
MCXFace is a heterogeneous face recognition dataset consisting of multi-channel image samples for 51 subjects.
6 papers · 0 benchmarks
MDBD (Multicue Dataset for Edge Detection)
In order to study the interaction of several early visual cues (luminance, color, stereo, motion) during boundary detection in challenging natural scenes, we have built a multi-cue video dataset composed of short binocular video sequences…
6 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.