Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 187 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 8929–8976 of 12,172
KITTI-6DoF is a dataset that contains annotations for the 6DoF estimation task for 5 object categories on 7,481 frames.
1 paper · 0 benchmarks
Please find more details of this dataset at https://alex-xun-xu.github.io/ProjectPage/CVPR18/index.html 3D motion segmentation has been the key problem in computer vision research due to the application in structure from motion and…
1 paper · 1 benchmark
KTI Multiview Football II consists of images of professional footballers during a match of the Allsvenskan league.
1 paper · 0 benchmarks
A new word analogy task dataset for Indonesian.
1 paper · 0 benchmarks
This dataset was build as a part of development of Treebanks for Indian Languages funded by MeitY, Govt.
1 paper · 0 benchmarks
At KayifamilyTv, we introduce a powerful and scalable video-mining solution designed to enhance captioning capabilities by transferring supervision from image datasets to video and audio content.
1 paper · 0 benchmarks
Context The Kepler Space Observatory is a NASA-build satellite that was launched in 2009.
1 paper · 1 benchmark
Dataset contains CS/Math articles abstracts (in Russian) obtained from two online sources.
1 paper · 0 benchmarks
Handwriting analysis is still an important application in machine learning.
1 paper · 0 benchmarks
Kinect-WSJ is a multichannel, multispeaker, reverberated, noisy dataset which extends the WSJ0-2mix singlechannel, non-reverberated, noiseless dataset to the strong reverberation and noise conditions and the Kinect-like microphone array…
1 paper · 0 benchmarks
Kinetics-GEB+ (Generic Event Boundary Captioning, Grounding and Retrieval) is a dataset that consists of over 170k boundaries associated with captions describing status changes in the generic events in 12K videos.
1 paper · 3 benchmarks
Explicitly created for Human Computer Interaction (HCI).
1 paper · 0 benchmarks
This dataset comprises over 9,000 images captured in the AI2-THOR simulation environment, featuring 69 distinct object classes.
1 paper · 0 benchmarks
The Kite database is a multi-modal dataset for the control of unmanned aerial vehicles (UAVs).
1 paper · 0 benchmarks
Knot128 is a dataset to test knot untangling algorithms, i.e., highly-tangled configurations that can be difficult to smooth out into a canonical knot embedding.
1 paper · 0 benchmarks
We introduce KnowledJe, an English-language knowledge graph of antisemitic history and language from the 20th century to the present.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Full Koo platform Dataset.
1 paper · 0 benchmarks
Kor-Lang8 is a Korean grammatical error correction (GEC) dataset extracted from the NAIST Lang-8 Learner Corpora by the language label.
1 paper · 0 benchmarks
Kor-Learner is a Korean grammatical error correction (GEC) dataset made from the NIKL learner corpus containing essays written by Korean learners and their grammatical error correction annotations by their tutors in an morpheme-level XML…
1 paper · 0 benchmarks
Kor-Learner is a Korean grammatical error correction (GEC) dataset collected grammatically from two sources, and the correct sentences were read using Google Text-to-Speech(TTS) system.
1 paper · 0 benchmarks
The data contains the following attributes for Korea Stock Price Index (KOSPI) for January 2000–December 2016: 1.
1 paper · 1 benchmark
APEACH is the first crowd-generated Korean evaluation dataset for hate speech detection.
1 paper · 0 benchmarks
1.9K Korean Online Hate Speech Comments for Multilabel Classification (Annotated by Three Independent Labelers per Data)
1 paper · 0 benchmarks
OHLCVT stands for Open, High, Low, Close, Volume and Trades and represents the following trading information within each time frame (such as one minute, five minute, hourly, daily, etc.): + Open: the first traded price + High: the highest…
1 paper · 0 benchmarks
a parallel corpus of Sorani (ckb or Central Kurdish) and Kurmanji (kmr or Northern Kurdish) dialects of Kurdish along with English (eng).
1 paper · 0 benchmarks
Cleaned and preprocessed version of the Kyokushin Karate Motion Dataset by Szczkesna et al.
1 paper · 0 benchmarks
L-CAS 3D Point Cloud People Dataset contains 28,002 Velodyne scan frames acquired in one of the main buildings (Minerva Building) of the University of Lincoln, UK.
1 paper · 0 benchmarks
The Sentinel-2 satellite carries 12 CMOS detectors for the VNIR bands, with adjacent detectors having overlapping fields of view that result in overlapping regions in level-1 B (L1B) images.
1 paper · 0 benchmarks
L3Cube-MahaCorpus is a Marathi monolingual data set scraped from different internet sources.
1 paper · 0 benchmarks
LADI v2 (Low Altitude Disaster Imagery v2)
LADI Overview The Low Altitude Disaster Imagery (LADI) dataset was created to address the relative lack of annotated post-disaster aerial imagery in the computer vision community.
1 paper · 0 benchmarks
LAMBDA (Long-term Ad MemoraBility DAtaset)
LAMDBA is a long term ad memorability dataset, featuring data from 1749 participants and 2205 ads across 276 brands.
1 paper · 0 benchmarks
Raw negotiation transcripts generated for the paper "Evaluating Language Model Agency through Negotiations".
1 paper · 0 benchmarks
An ancestral origin database of 14,000 images of individuals from East Asia, the Indian subcontinent, sub-Saharan Africa, and Western Europe.
1 paper · 0 benchmarks
LARQS (An Evaluation Dataset for Chinese Codex Word Embedding Model)
Word embedding is a modern distributed word representations approach widely used in many natural language processing tasks.
1 paper · 0 benchmarks
LARa (Logistic Activity Recognition Challenge)
LARa is the first freely accessible logistics-dataset for human activity recognition.
1 paper · 0 benchmarks
Large Shape and Texture dataset (LAS&T) is a giant dataset of shapes and textures for tasks of visual shapes and textures identification and retrieval from single image.
1 paper · 0 benchmarks
LAVIB (Large-scale Video Interpolation Benchmark)
LAVIB comprises a large collection of high-resolution videos sourced from the web.
1 paper · 1 benchmark
Cosmic rays in the LCO CR dataset are labeled accurately and consistently across many diverse observations from various instruments.
1 paper · 0 benchmarks
For our experiments, we collected a dataset of procedural knowledge of the LangChain Python library, unseen by many extant LLMs.
1 paper · 0 benchmarks
LDCT-and-Projection-data 医学去噪数据集
1 paper · 0 benchmarks
LDD (LDD: A Grape Diseases Dataset Detection and Instance Segmentation)
The Instance Segmentation task, an extension of the well-known Object Detection task, is of great help in many areas, such as precision agriculture: being able to automatically identify plant organs and the possible diseases associated…
1 paper · 2 benchmarks
LDDRS (LWIR DoFP Dataset of Road Scene)
The LWIR DoFP Dataset of Road Scene (LDDRS) is a road detection dataset with 2,113 annotated images.
1 paper · 0 benchmarks
The datasets of "Towards Lightweight Cross-domain Sequential Recommendation via External Attention-enhanced Graph Convolution Network" (DASFAA 2023)
1 paper · 0 benchmarks
The dataset was collected from two courses offered on the University of Jordan's E-learning Portal during the second semester of 2020, namely "Computer Skills for Humanities Students" (CSHS) and "Computer Skills for Medical Students"…
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.