Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 217 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 10369–10416 of 12,172
The StudyAbroadGPT-Dataset is a collection of conversational data focused on university application requirements for various programs, including MBA, MS in Computer Science, Data Science, and Bachelor of Medicine.
1 paper · 0 benchmarks
Digital Edition: Sturm Edition Source: Schrade, Torsten: „Startseite“, in: DER STURM.
1 paper · 0 benchmarks
The SuSy Dataset combines authentic photographs and AI-generated images designed for training and evaluating synthetic image detection models.
1 paper · 0 benchmarks
SubSumE Dataset This repository contains the SubSumE dataset for subjective document summarization.
1 paper · 0 benchmarks
This repository contains replication data to the paper titled: "Anti-noise window: subjective perception of active noise reduction and effect of informational masking"
1 paper · 0 benchmarks
A large dataset of around 40000 Reddit posts was collected from r/suicidewatch and other non-suicidal subreddits.
1 paper · 0 benchmarks
The dataset contains 140 paragraphs from climate change reports with associated aspect-based (i.e.
1 paper · 0 benchmarks
SumeCzech-NER contains named entity annotations of SumeCzech 1.0, a Czech news-based summarization dataset.
1 paper · 0 benchmarks
SummZoo, a benchmark consists of 8 diverse summarization tasks with multiple sets of few-shot samples for each task, covering both monologue and dialogue domains.
1 paper · 0 benchmarks
SunspotsYoloDataset is a set of 1690+380+128 high-resolution RGB astronomical images captured with smart telescopes with specific solar filters and annotated with the positions of sunspots that are effectively in the images.
1 paper · 0 benchmarks
Super-CLEVR-3D is a visual question answering (VQA) dataset where the questions are about the explicit 3D configuration of the objects from images (i.e.
1 paper · 0 benchmarks
The Super-resolution of Multi-Dimensional Diffusion MRI (Super MUDI) dataset contains the data of four healthyhuman subjects with ages range between 19 and 46 years.
1 paper · 0 benchmarks
We introduce SuperRS-VQA (avg.
1 paper · 0 benchmarks
This deposit is supplementary material to "Machine Learning Applications in Archaeological Practices: A Review".
1 paper · 0 benchmarks
The file contains an annotated list of papers that are included in the literature survey.
1 paper · 0 benchmarks
Accompanying supplementary data for the paper.
1 paper · 0 benchmarks
The SurfaceGrid dataset contains nearly a million 512x512 images for use in training neural networks on shape-fron-surface contour task.
1 paper · 0 benchmarks
Dataset Generation - Base Model: h2oai/h2ogpt-gm-oasst1-en-2048-falcon-40b-v2 - Seed Instructions: Selected from the databricks/databricks-dolly-15k dataset - Generation Approach: Iterative evolution of instructions using a conversational…
1 paper · 0 benchmarks
Overview The LaMini Dataset is an instruction dataset generated using h2ogpt-gm-oasst1-en-2048-falcon-40b-v2.
1 paper · 0 benchmarks
Dataset Generation - Base Model: h2oai/h2ogpt-gm-oasst1-en-2048-falcon-40b-v2 - Seed Instructions: Derived from the FLAN-v2 Collection.
1 paper · 0 benchmarks
Surgical Hands is a dataset that provides multi-instance articulated hand pose annotations for in-vivo videos.
1 paper · 0 benchmarks
The dataset is collected from the Youtube videos that contains fight instances in it.
1 paper · 0 benchmarks
Survey answers (Answers to surveys in both papers, as well as processed answers)
Please see paper for questions.
1 paper · 0 benchmarks
Annotating data is a time-consuming and costly task, but it is inherently required for supervised machine learning.
1 paper · 0 benchmarks
There are 9,321 survey papers with high quality included in the SurvayBank in the domain of computer science.
1 paper · 0 benchmarks
The dataset contains cardiovascular medical records taken from 299 patients.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
A corpus of annotated crime stories from English-language newspapers in the U.S.
1 paper · 0 benchmarks
To explore the nascent area of sustainable venture capital, a review of related research was conducted and social entrepreneurs & investors interviewed to construct a questionnaire assessing the interests and intentions of current & future…
1 paper · 0 benchmarks
A dataset of images containing leaves from 15 tree classes.
1 paper · 0 benchmarks
Switchboard Dialog Act Corpus
1 paper · 1 benchmark
SyDog (A Synthetic Dog Dataset)
SyDog is a synthetic dataset of dogs containing ground truth pose and bounding box coordinates which was generated using the game engine, Unity.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This is a pose estimation dataset, consisting of symmetric 3D shapes where multiple orientations are visually indistinguishable.
1 paper · 0 benchmarks
Symmetry-OOD is a dataset for symmetry perception by deep neural networks.
1 paper · 0 benchmarks
Table extraction from document images is a challenging AI problem, and labelled data for many content domains is difficult to come by.
1 paper · 0 benchmarks
50K synthetic renders of the human foot, with surface normals, masks and keypoints.
1 paper · 0 benchmarks
SynoClip Dataset The SynoClip dataset is a comprehensive and standard dataset specifically designed for the video synopsis task.
1 paper · 0 benchmarks
3D Computer Graphics is leveraged to generate a large and diverse dataset for training bike rotation estimators in bike parking assessment.
1 paper · 0 benchmarks
A dataset consisting of high-quality, synthetic chest X-rays from the CheXGenBench-benchmark leading model, Sana (0.6B).
1 paper · 0 benchmarks
A public open dataset of synthetic chest X-ray images of COVID-19.
1 paper · 0 benchmarks
The Synthetic COVID-19 Chest X-ray Dataset consists of 21,295 synthetic COVID-19 chest X-ray images to be used for computer-aided diagnosis.
1 paper · 0 benchmarks
This dataset accompanies the paper Learning the mechanisms of network growth' by the same authors.
1 paper · 1 benchmark
This is the first federated quantum dataset in the literature.
1 paper · 0 benchmarks
A synthetic dataset for evaluating non-rigid 3D human reconstruction based on conventional RGB-D cameras.
1 paper · 0 benchmarks
This dataset is a large-scale synthetic dataset to simulate the attack scenario for a keystroke inference attack.
1 paper · 0 benchmarks
This dataset is originally created for the Knowledge Graph Reasoning Challenge for Social Issues (KGRC4SI) Video data that simulates daily life actions in a virtual space from Scenario Data.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.