Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 186 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 8881–8928 of 12,172
Iran's Built Heritage Binary Image Classification Dataset contains approximately 10,500 CHB images gathered from four different sources: i) The archives of Iran’s cultural heritage ministry ii) The author’s (M.B) personal archives iii)…
1 paper · 0 benchmarks
The Iranis Dataset is a Large-scale dataset of Farsi license plate characters containing a large-scale dataset with more than 83,000 images of Farsi numbers and letters collected from real-world license plate images captured by various…
1 paper · 0 benchmarks
Labelled dataset of Iridium “ring alert” downlink messages, including message headers captured at 25MS/s.
1 paper · 0 benchmarks
Text from Irish Wikipedia, an online encyclopedia.
1 paper · 0 benchmarks
The Istella LETOR full dataset is composed of 33,018 queries and 220 features representing each query-document pair.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
JAAH (Jazz Audio-Aligned Harmony)
Eremenko, E.
1 paper · 2 benchmarks
JAMBO (A Multi-Annotator Image Dataset for Benthic Habitat Classification)
The JAMBO dataset contains 3290 underwater images of the seabed captured by an ROV in temperate waters in the Jammer Bay area off the North West coast of Jutland, Denmark.
1 paper · 0 benchmarks
A corpus of about 115 hours of Dutch speech from juveniles, non-native speakers and seniors, consisting of read text and man-machine dialogues.
1 paper · 0 benchmarks
JAZZVAR Dataset (JAZZVAR: A Dataset of Variations found within Solo Piano Performances of Jazz Standards for Music Overpainting)
Jazz pianists often uniquely interpret jazz standards.
1 paper · 0 benchmarks
JDDC 2.0 is a large-scale multimodal multi-turn dialogue dataset collected from a mainstream Chinese E-commerce platform JD.com, containing about 246 thousand dialogue sessions, 3 million utterances, and 507 thousand images, along with…
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
JECC (Jericho Environment Commonsense Comprehension)
Jericho Environment Commonsense Comprehension (JECC) is a dataset for commonsense reasoning.
1 paper · 0 benchmarks
Involves data where a robot interacts with 5.1 cm colored blocks to complete an order-fulfillment style block stacking task.
1 paper · 0 benchmarks
The JNU Bearing Dataset, developed by Jiangnan University in China, is widely used in the field of fault diagnosis for rotating machinery.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The JPersonaChat dataset is built by NTT CS LAB for Japanese dialog transformers models.
1 paper · 0 benchmarks
We processed 241 pairs of CXR and DES soft tissue images from the JSRT dataset by performing operations like inversion and contrast adjustment to convert these images into negative formats more frequently used in clinical settings.
1 paper · 0 benchmarks
The Jejueo Single Speaker Speech (JSS) dataset consists of 10k high-quality audio files recorded by a native Jejueo speaker and a transcript file.
1 paper · 0 benchmarks
The information contained in JUSThink Dialogue and Actions Corpus dataset includes dialogue transcripts, event logs, and test responses of children aged 9 through 12, as they participate in a robot-mediated human-human collaborative…
1 paper · 0 benchmarks
Jacquard v2 (Jacquard V2: Refining Datasets using the Human In the Loop Data Correction Method)
In the context of rapid advancements in industrial automation, vision-based robotic grasping plays an increasingly crucial role.
1 paper · 0 benchmarks
JamPatoisNLI provides the first dataset for natural language inference in a creole language, Jamaican Patois.
1 paper · 1 benchmark
📊 Dataset Details - Name: JamendoMaxCaps - URL: https://huggingface.co/datasets/amaai-lab/JamendoMaxCaps - Content: 362,238 songs with captions generated by Qwen2-Audio Metadata Fields - genre - speed - variable tags 🎯 Rationale 1.
1 paper · 0 benchmarks
Click to #Development of QA-SQP for non-linear and history-dependent mechanical problems This directory contains the source code and numerical benchmarks published in [^1] Dependencies and Prerequisites Python, pandas, numpy, matplotlib…
1 paper · 0 benchmarks
The Java dataset introduced in Hybrid-DeepCom (Deep code comment generation with hybrid lexical and syntactical information), commonly used to evaluate automated code summarization.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
JoCAD is a dataset for anomaly detection in citation networks.
1 paper · 0 benchmarks
All possible intermediate result cardinalities for 3300 queries on IMDb.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Source: Text mining methodologies with R: An application to central bank texts
1 paper · 0 benchmarks
Diffusion generated image dataset
1 paper · 0 benchmarks
A high-definition Talking Face dataset featuring Chinese-language videos - 1.1k videos, 130h - high-resolution video frames, average 400+ pixels of detected faces - durations ranging from 46 seconds to 52 minutes - approximately equal…
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
JustLogic is a natural language deductive reasoning dataset.
1 paper · 0 benchmarks
Korean Multi-label Hate Speech Dataset We introduce K-MHaS, a new multi-label dataset for hate speech detection that effectively handles Korean language patterns.
1 paper · 0 benchmarks
K-QA(fa) (persian translation of K-QA dataset)
persian translation of K-QA dataset
1 paper · 0 benchmarks
The KACC benchmark consists of three subtasks that can be applied to knowledge graphs: knowledge abstraction, knowledge concretization and knowledge completion.
1 paper · 0 benchmarks
This is the dataset for testing the robustness of various VO/VIO methods, acquired on reak UAV.
1 paper · 0 benchmarks
KAgentBench is a benchmark dataset of over 3,000 human-edited, automated evaluation data for testing agent capabilities, with evaluation dimensions including planning, tool use, reflection, concluding, and profiling.
1 paper · 0 benchmarks
KArSL (Arabic Sign Language Dataset)
KArSL (KFUPM Arabic Sign Language) is an Arabic sign language (ArSL) database collected using Microsoft Kinect V2.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
KCIF (Knowledge Conditioned Instruction Following (KCIF))
KCIF is a benchmark for evaluating the instruction-following capabilities of Large Language Models (LLM).
1 paper · 0 benchmarks
KD-EmoR (Korean Drama Scene Transcript Dataset for Emotion Recognition in Conversations)
KD-EmoR is socio-behavioral emotion dataset for emotion recognition in realistic conversation scenarios.
1 paper · 1 benchmark
KGRED (Knowledge-graph-enhanced relation extraction datasets--)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
KHATT (KFUPM Handwritten Arabic TexT Database)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Two news datasets (KINNEWS and KIRNEWS) for multi-class classification of news articles in Kinyarwanda and Kirundi, two low-resource African languages.
1 paper · 0 benchmarks
Extension of the official KITTI'15 dataset.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.