Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 150 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 7153–7200 of 12,172
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
This dataset was collected during a LoRaWAN measurement campaign in a multi-room indoor office environment in the University of Siegen, Germany.
2 papers · 0 benchmarks
iBugMask is an in-the-wild face parsing dataset that contains 1,000 challenging face images and manually annotated labels for 11 semantic classes: background, facial skin, left/right brow, left/right eye, nose, upper/lower lip, inner…
2 papers · 1 benchmark
iFF (Intrinsic Forward Facing)
Real-world dataset on forward facing scenes with different camera intrinisc parameters.
2 papers · 1 benchmark
A dataset for fine-grained art attribute recognition introduced in the 6th FGVC Workshop at CVPR 2019.
2 papers · 0 benchmarks
iRodent (iRodent Animal Pose Estimation)
Description: The "iRodent" dataset contains rodent species observations obtained using the iNaturalist API, with a focus on Suborder Myomorpha (Taxon ID: 16).
2 papers · 1 benchmark
The dataset is an enhanced version of the im2latex-100k dataset.
2 papers · 0 benchmarks
ivrit.ai (database of Hebrew audio and text content.)
ivrit.ai is a database of Hebrew audio and text content.
2 papers · 0 benchmarks
kickstarter (Funding Successful Projects on Kickstarter)
Kickstarter is a community of more than 10 million people comprising of creative, tech enthusiasts who help in bringing creative project to life.
2 papers · 1 benchmark
KITAB is a challenging dataset and a dynamic data collection approach for testing abilities of Large Language Models (LLMs) in answering information retrieval queries with constraint filters.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
Large language models such as ChatGPT and GPT-4 have recently achieved astonishing performance on a variety of natural language processing tasks.
2 papers · 0 benchmarks
A data set of Sudoku grids with more than one solution.
2 papers · 0 benchmarks
mini-ImageNet was proposed by Matching networks for one-shot learning for few-shot learning evaluation, in an attempt to have a dataset like ImageNet while requiring fewer resources.
2 papers · 1 benchmark
Overview nEMO is a simulated dataset of emotional speech in the Polish language.
2 papers · 0 benchmarks
neuronIO (Single cortical neuron (L5PC) input output simulation at 1ms temporal resolution)
Single cortical neurons as deep artificial neural networks This dataset contains training and testing subsets of the input/output relationship of a single cortical layer 5 pyramidal cell (L5PC) neuron at 1ms single spike temporal…
2 papers · 0 benchmarks
news20 (NewsWeeder: learning to filter netnews)
Two datasets featuring binary and multi-class classification.
2 papers · 0 benchmarks
We introduce a novel dataset consisting of images depicting pink eggs that have been identified as Pomacea canaliculata eggs, accompanied by corresponding bounding box annotations.
2 papers · 0 benchmarks
The pioNER corpus provides gold-standard and automatically generated named-entity datasets for the Armenian language.
2 papers · 0 benchmarks
robo-vln (Robotics Vision-and-Language Navigation)
The Robo-VLN dataset is a continuous control formulation of the VLN-CE dataset by Krantz et al ported over from Room-to-Room (R2R) dataset created by Anderson et al.
2 papers · 1 benchmark
A set of 180,000 Sudoku grids with a variable number of hints from the minimal number of 17 (extremely hard instances) to 34 (easy instances), with 10,000 instances per level of hardness.
2 papers · 0 benchmarks
A set of easy Sudoku instances used in the SATNet paper for training SatNet on how to learn to play Sudoku.
2 papers · 0 benchmarks
satp-zsm-stage1 (Replication Data for: Crossing the Linguistic Causeway: A Binational Approach for Translating Soundscape Attributes to zsm)
This is the replication data for the paper: "Crossing the Linguistic Causeway: A Binational Approach for Translating Soundscape Attributes to Bahasa Melayu".
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
Clean version of UDHR (Universal Declaration of Human Rights), at the long sentence level.
2 papers · 0 benchmarks
A total of 18 sequences were collected of various lengths.
2 papers · 0 benchmarks
wifi_data (WiFi Data for HMM Anomaly Detection)
Wi-Fi dataset: the dataset may be downloaded from this link.
2 papers · 0 benchmarks
Unsustainable fishing practices worldwide pose a major threat to marine resources and ecosystems.
2 papers · 1 benchmark
The data used for !Optimizer 2021 competition, based on seven biological model organisms.
1 paper · 0 benchmarks
chinahate dataset contains a total of 2,172,333 tweets hashtagged #china posted during the time it was collected.
1 paper · 0 benchmarks
10kGNAD (Ten Thousand German News Articles Dataset)
The 10kGNAD dataset is intended to solve part of this problem as the first German topic classification dataset.
1 paper · 0 benchmarks
Undress AI apps, powered by advanced AI and deep learning, have sparked both curiosity and controversy.
1 paper · 0 benchmarks
In one round of sequencing, 5 fecal pellets from 2 pro-inflammatory environments (Harvard BRI/Johns Hopkins) and 2 pro-survival environments (Broad Institute/Jackson Labs) were sequenced at the 16s rDNA locus.
1 paper · 0 benchmarks
The 1DSfM Landmarks is a collection of community-based image reconstruction by Kyle Wilson and is comprised of 14 datasets with comparison to bundler ground truth.
1 paper · 0 benchmarks
The 2007 BP TTI Velocity-Analysis Benchmark dataset was created by Hemang Shah and provided courtesy of BP Exploration Operation Company Limited ("BP").
1 paper · 0 benchmarks
Description: 23 Pairs of Identical Twins Face Image Data.
1 paper · 0 benchmarks
Dataset contains images of dogs and cats.
1 paper · 0 benchmarks
Official dataset for Towards A Holistic Landscape of Situated Theory of Mind in Large Language Models.
1 paper · 0 benchmarks
Dataset of low fidelity resolutions of the RANS equations over airfoils.
1 paper · 0 benchmarks
We provide here a new multi-view text dataset, collected from three well-known online news sources: BBC, Reuters, and The Guardian.
1 paper · 0 benchmarks
Dataset for the 32 years of IEEE VIS
1 paper · 0 benchmarks
The 360+x dataset is a large-scale database that emphasizes a comprehensive multifaceted understanding of daily scenes.
1 paper · 0 benchmarks
The dataset consists of high-resolution three-dimensional (3D) turbulent flow simulations.
1 paper · 0 benchmarks
3D design file repository for the Stickbug Robot a 6 armed holonomic precision pollination robot
1 paper · 0 benchmarks
3D-BSLS-6D (3D scans of Bins by Structured-Light Scanner for 6D pose estimation)
Dataset consist of both real captures from Photoneo PhoXi structured light scanner devices annotated by hand and synthetic samples produced by custom generator.
1 paper · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.