Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 234 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 11185–11232 of 12,172
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
distillation and psychological sft using only one dataset data statement psychological knowledge : general knowledge ≈ 4 : 10 (precise num is 3868 : 10000) > In jsonl file, the first 10k lines refer to general knowledge while the rest…
1 paper · 0 benchmarks
The publicmeetings corpus contains meetings, made of pairs of automatic transcriptions from audio recordings and meeting reports written by a professional.
1 paper · 0 benchmarks
pursuitMW (Multi-agent pursuit in matrix world)
Multi-agent pursuit in matrix world (pursuitMW) is a partially observable Markov game (POMG) between a swarm of pursuers and a swarm of evaders.
1 paper · 0 benchmarks
This dataset is made of 6366 threads collected from the r/AmITheAsshole community on Reddit.
1 paper · 0 benchmarks
Questions regarding computer science education for members of the r/transprogrammer Reddit.
1 paper · 0 benchmarks
Data was collected from Tobii Fusion screen-based Eye Tracker.
1 paper · 0 benchmarks
rc_49 (rc_49 Grasping Dataset)
Includes several sets of synthetic stereo images labelled with grasp rectangles representing parallel-jaw grasps (Cornell-like format).
1 paper · 0 benchmarks
Reader eye tracking and engagement scores for two short stories, aggregated by sentence.
1 paper · 0 benchmarks
This the dataset for Every Language Counts: Learn and Unlearn in Multilingual LLMs.
1 paper · 0 benchmarks
County-level and municipal-level data on private election administration grant receipt, census data, and election outcomes all in tabular form
1 paper · 0 benchmarks
Repository of containerized services that can be migrated through UMS.
1 paper · 0 benchmarks
The results-A dataset is a dataset consisting of 22 infrared images commonly used for testing performance of Infrared Image Super-Resolution models.
1 paper · 1 benchmark
The results-C dataset is a dataset consisting of 22 infrared images commonly used for testing performance of Infrared Image Super-Resolution models.
1 paper · 1 benchmark
The dataset contains 71 samples with (normalized) expression data for 4,088 genes.
1 paper · 0 benchmarks
Simulations and hardware experiments of different coverage tasks saved as robot swarm objects.
1 paper · 0 benchmarks
The Innodata Red Teaming Prompts aims to rigorously assess models’ factuality and safety.
1 paper · 1 benchmark
The Innodata Red Teaming Prompts aims to rigorously assess models’ factuality and safety.
1 paper · 0 benchmarks
The Innodata Red Teaming Prompts aims to rigorously assess models’ factuality and safety.
1 paper · 1 benchmark
sPBC (Super Parallel Bible Corpus)
This is a super-parallel Bible corpus containing 1401 language labels (languagescript pairs), meaning that for each verse, we have the translation in other languages.
1 paper · 0 benchmarks
The sRGB2XYZ dataset contains ~1,200 pairs of camera-rendered sRGB and the corresponding scene-referred CIE XYZ images (971 training, 50 validation, and 244 testing images).
1 paper · 0 benchmarks
satp-zsm-stage2 (Replication data for Crossing the Linguistic Causeway: Ethnonational Differences on Soundscape Attributes in Bahasa Melayu)
This is the replication data for the paper: "Crossing the Linguistic Causeway: Ethnonational Differences on Soundscape Attributes in Bahasa Melayu".
1 paper · 0 benchmarks
scb-mt-en-th-2020 is an English-Thai machine translation dataset with over 1 million segment pairs, curated from various sources, namely news, Wikipedia articles, SMS messages, task-based dialogs, web-crawled data and government documents.
1 paper · 0 benchmarks
Appendix A in this paper contains a real-world name length data for the whole of Sweden as well as Stockholm Municipality (Swedish: Stockholms kommun) as of 31 December 2019.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Data simulator for polypharmacies / drug combinations TL;DR python createdataset.py [--config path/to/config.json --seed yourseed] Template of config.json in configs/ Example end result | Rx1 | Rx2 | Rx3 | Rx4 | ...
1 paper · 0 benchmarks
These datasets, ComCo and SimCo, designed for evaluating multi-object representation in Vision-Language Models (VLMs).
1 paper · 0 benchmarks
The simply-CLEVR dataset aims to provide a benchmark dataset that can be used for transparent quantitative evaluation of explanation methods (aka heatmaps/XAI methods).
1 paper · 0 benchmarks
sonar (Connectionist Bench (Sonar, Mines vs. Rocks))
The task is to train a network to discriminate between sonar signals bounced off a metal cylinder and those bounced off a roughly cylindrical rock.
1 paper · 0 benchmarks
The dataset contains standard contexts of the lattices of all atomic lattices in the Concept Explorer format.
1 paper · 0 benchmarks
The Stickerchat dataset is a large-scale real-world dialog dataset with stickers which contains 340K multi-turn dialog and sticker pairs.
1 paper · 0 benchmarks
This collection contains 156 cases of MAASTRO Head and Neck images and RTStruct contours.
1 paper · 0 benchmarks
The synRailObs contains following categories: - Person - Rocks - Vehicles - Moto-cars - Animals Each directory contain images and correspoing yolo-format annotations and masks, which can be leveraged in both object detection and…
1 paper · 0 benchmarks
The synethetic dataset (10000 pairs of images and region, 2.95GB) is shared with the code (hdf5 dataset format).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This dataset is a patched version of The Taste & Affect Music Database by D.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The datasets of "Time Interval-enhanced Graph Neural Network for Shared-account Cross-domain Sequential Recommendation" (TNNLs 2022)
1 paper · 0 benchmarks
This dataset contains data scraped from search results of the query #israel and #palestine during the early months of the 2023 crisis.
1 paper · 0 benchmarks
This deposit contains benchmark code, data and results to assess the Python software time-agnostic-library v4.3.0.
1 paper · 0 benchmarks
titanic5 Dataset Created by David Beltran del Rio March 2016.
1 paper · 0 benchmarks
topex-printer is a dataset containing 102 machine parts of a label printing machine.
1 paper · 0 benchmarks
Dataset based on Twitter usernames of American politicians.
1 paper · 0 benchmarks
Microscopy is a cornerstone of biomedical research, enabling detailed study of biological structures at multiple scales.
1 paper · 0 benchmarks
The uniD dataset is an innovative collection of naturalistic road user trajectories, captured within the RWTH Aachen University campus using drone technology to address common challenges such as occlusions found in traditional traffic data…
1 paper · 0 benchmarks
This dataset contains the ground truth for urban changes occurred in Mariupol, Ukraine for the time frame 2017-2020.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.