Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 114 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 5425–5472 of 12,172
HelixNet (HelixNet: A Dataset for Online LiDAR Segmentation)
Large-scale and open-access LiDAR dataset intended for the evaluation of real-time semantic segmentation algorithms.
3 papers · 1 benchmark
Hephaestus (Hephaestus: A large scale multitask dataset towards InSAR understanding)
Hephaestus is the first large-scale InSAR dataset.
3 papers · 0 benchmarks
HiRID is a freely accessible critical care dataset containing data relating to almost 34 thousand patient admissions to the Department of Intensive Care Medicine of the Bern University Hospital, Switzerland (ICU), an interdisciplinary…
3 papers · 6 benchmarks
dataset link : https://www.kaggle.com/datasets/osamahosamabdellatif/high-quality-invoice-images-for-ocr Overview High-Quality Invoice Images for OCR is a curated dataset containing professionally scanned and digitally captured invoice…
3 papers · 0 benchmarks
A dataset for benchmarking action recognition algorithms in natural environments, while making use of 3D information.
3 papers · 0 benchmarks
Publicly available dataset in the hotel domain (50M versus 0.9M) and additionally, the largest recommendation dataset in a single domain and with textual reviews (50M versus 22M).
3 papers · 0 benchmarks
HuRDL (Human-Robot Dialogue Learning Corpus)
The Human-Robot Dialogue Learning (HuRDL) Corpus is a dataset about asking questions in situated task-based interactions.
3 papers · 0 benchmarks
A synthetic data of videos of human action sequences and the corresponding optical flow.
3 papers · 0 benchmarks
The Human Protein Atlas contains images of histological sections from normal and cancer tissues obtained by immunohistochemistry.
3 papers · 3 benchmarks
We present datasets containing urban traffic and rural road scenes recorded using hyperspectral snap-shot sensors mounted on a moving car.
3 papers · 1 benchmark
Hyperpartisan News Detection was a dataset created for PAN @ SemEval 2019 Task 4.
3 papers · 1 benchmark
IAM Dataset (A Comprehensive and Large-Scale Dataset for Integrated Argument Mining Tasks)
We introduce a large and comprehensive dataset to facilitate the study of several essential AM tasks in the debating system.
3 papers · 2 benchmarks
IBC (Individual Brain Charting)
The Individual Brain Charting (IBC) project aims at providing a new generation of functional-brain atlases.
3 papers · 0 benchmarks
ICSI Meeting Corpus in JSON format.
3 papers · 1 benchmark
The ICVL dataset is a hand pose estimation dataset that consists of 330K training frames and 2 testing sequences with each 800 frames.
3 papers · 0 benchmarks
Please refer: https://github.com/google/imageinwords/blob/main/datasets/IIW-400/README.md
3 papers · 0 benchmarks
IJB-S (IARPA Janus Benchmark-S)
Paper Abstract We present IJB–S dataset, an open-source IARPA Janus Surveillance Video Benchmark and associated protocols.
3 papers · 1 benchmark
We have cleaned the noisy IMDB-WIKI dataset using a constrained clustering method, resulting this new benchmark for in-the-wild age estimation.
3 papers · 1 benchmark
IPAC (Icelandic Parallel Abstracts Corpus)
IPAC (Icelandic Parallel Abstracts Corpus ) is a new Icelandic-English parallel corpus, composed of abstracts from student theses and dissertations.
3 papers · 0 benchmarks
This dataset contains the data for the paper 'Using Multiple Instance Learning for Explainable Solar Flare Prediction'.
3 papers · 0 benchmarks
In The Groove (ITG) is an audio dataset where given a raw audio track, the goal is to produce a choreography step chart, similar to those used in the Dance Dance Revolution video game.
3 papers · 0 benchmarks
The IWSLT 2019 dataset contains source, Machine Translated, reference and Post-Edited text, which can be used to quantify and evaluate Post-editing effort after automatic MT.
3 papers · 0 benchmarks
IfAct (Identifying Human Actions Visible in Online Vlogs)
We consider the task of identifying human actions visible in online videos.
3 papers · 0 benchmarks
IllusionVQA is a Visual Question Answering (VQA) dataset with two sub-tasks.
3 papers · 2 benchmarks
A dataset of description sequences, a sequence of expressions that together are meant to single out one image from an (imagined) set of other similar images.
3 papers · 0 benchmarks
The Image and Video Advertisements collection consists of an image dataset of 64,832 image ads, and a video dataset of 3,477 ads.
3 papers · 0 benchmarks
We introduce diverse and realistic backgrounds into the images or color, texture, and adversarial changes in the background
3 papers · 0 benchmarks
transform the ImageNet-1K classification datatset for Chinese models by translating labels and prompts into Chinese.
3 papers · 1 benchmark
Imgur5k is a large-scale handwritten in-the-wild dataset, containing challenging real world handwritten samples from nearly 5K writers.
3 papers · 0 benchmarks
The IndicNLP corpus is a large-scale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families.
3 papers · 0 benchmarks
A special corpus of Indian languages covering 13 major languages of India.
3 papers · 13 benchmarks
The Indoor-6 dataset was created from multiple sessions captured in six indoor scenes over multiple days.
3 papers · 0 benchmarks
InfantMarmosetsVox is a dataset for multi-class call-type and caller identification.
3 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 0 benchmarks
Itihasa is a large-scale corpus for Sanskrit to English translation containing 93,000 pairs of Sanskrit shlokas and their English translations.
3 papers · 1 benchmark
JDsearch is a personalized product search dataset comprised of real user queries and diverse user-product interaction types (clicking, adding to cart, following, and purchasing) collected from JD.com, a popular Chinese online shopping…
3 papers · 0 benchmarks
A MIDI dataset of 500 4-part chorales generated by the KSChorus algorithm, annotated with results from hundreds of listening test participants, with 500 further unannotated chorales.
3 papers · 0 benchmarks
JUSTICE (JUSTICE: A Dataset for Supreme Court’s Judgment Prediction)
The dataset contains 3304 cases from the Supreme Court of the United States from 1955 to 2021.
3 papers · 0 benchmarks
JaNLI (Japanese Adversarial Natural Language Inference)
The Japanese Adversarial NLI (JaNLI) dataset is designed to require understanding of Japanese linguistic phenomena and illuminate the vulnerabilities of models.
3 papers · 0 benchmarks
The Jamendo Corpus is a voice detection dataset consisting of 93 songs with Creative Commons license from the Jamendo free music sharing website.
3 papers · 0 benchmarks
Includes 100K depth images under challenging scenarios.
3 papers · 2 benchmarks
KIEval provides a robust framework for dynamic, interactive evaluation of large language models, reducing the impact of data contamination and offering deeper insights into a model's true capabilities.
3 papers · 0 benchmarks
KIND (Kessler Italian Named-entities Dataset)
KIND is an Italian dataset for Named-Entity Recognition.
3 papers · 0 benchmarks
This Dataset consists of 2120 sequences of binary masks of pedestrians.
3 papers · 1 benchmark
KITTI is a well established dataset in the computer vision community.
3 papers · 0 benchmarks
KMIR (Knowledge Memorization, Identification, and Reasoning)
KMIR (Knowledge Memorization, Identification, and Reasoning) is a benchmark that covers 3 types of knowledge, including general knowledge, domain-specific knowledge, and commonsense, and provides 184,348 well-designed questions.
3 papers · 0 benchmarks
KOHTD (Kazakh Offline Handwritten Text Dataset)
Kazakh offline Handwritten Text dataset (KOHTD) has 3000 handwritten exam papers and more than 140335 segmented images and there are approximately 922010 symbols.
3 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.