Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 93 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 4417–4464 of 12,172
Open6DOR V2 (Benchmarking Open-instruction 6-DoF Object Rearrangement and A VLM-based Approach)
We introduce a challenging and comprehensive benchmark for open-instruction 6-DoF object rearrangement tasks, termed Open6DOR.
5 papers · 1 benchmark
A small RDF Knowledge Graph using FOAF and VCard.
5 papers · 0 benchmarks
Provides consistent ID annotations across multiple days, making it suitable for the extremely challenging problem of person search, i.e., where no clothing information can be reliably used.
5 papers · 0 benchmarks
PECAN (Paratope-Epitope Complexes for Antibody Networks (PECAN))
The PECAN dataset provides structural data for antibody-antigen interactions, specifically curated for paratope and epitope binding site prediction.
5 papers · 1 benchmark
PEN (Problems with Explanations for Numbers)
Provided explanations on the existing three benchmark datasets on solving algebraic word problems: ALG514, DRAW-1K, MAWPS
5 papers · 1 benchmark
PGDP5K (Plane Geometry Diagram Parsing Dataset)
PGDP5K is a dataset consisting of 5000 diagram samples composed of 16 shapes, covering 5 positional relations, 22 symbol types and 6 text types, labeled with more fine-grained annotations at primitive level, including primitive classes,…
5 papers · 1 benchmark
Please refer to the following paper which includes a description of the dataset and a link to the dataset and the paper code: Alain Hennebelle, Huned Materwal, and Leila Ismail, "HealthEdge: A Machine Learning-Based Smart Healthcare…
5 papers · 0 benchmarks
Year after year, the demand for ever-better smartphone photos continues to grow, in particular in the domain of portrait photography.
5 papers · 1 benchmark
This dataset contains 114 individuals including 1824 images captured from two disjoint camera views.
5 papers · 1 benchmark
PQuAD (Persian Question Answering Dataset)
Persian Question Answering Dataset (PQuAD) is a crowdsourced reading comprehension dataset on Persian Wikipedia articles.
5 papers · 0 benchmarks
A dataset composed of 12 different 3D scenes and RGB sequences of 20 subjects moving in and interacting with the scenes.
5 papers · 1 benchmark
PSI (IUPUI-CSRC Pedestrian Situated Intent)
The IUPUI-CSRC Pedestrian Situated Intent (PSI) benchmark dataset has two innovative labels besides comprehensive computer vision annotations.
5 papers · 0 benchmarks
ParaQA is a question answering (QA) dataset with multiple paraphrased responses for single-turn conversation over knowledge graphs (KG).
5 papers · 0 benchmarks
PatternNet is a large-scale high-resolution remote sensing dataset collected for remote sensing image retrieval.
5 papers · 1 benchmark
PedX is a large-scale multi-modal collection of pedestrians at complex urban intersections.
5 papers · 0 benchmarks
PerSenT is a dataset of crowd-sourced annotations of the sentiment expressed by the authors towards the main entities in news articles.
5 papers · 0 benchmarks
PoMo consists of more than 231K sentences with post-modifiers and associated facts extracted from Wikidata for around 57K unique entities.
5 papers · 0 benchmarks
The PodcastFillers dataset consists of 199 full-length podcast episodes in English with manually annotated filler words and automatically generated transcripts.
5 papers · 1 benchmark
Contains 10,000 fine-grained SKU-level products frequently bought by online customers in JD.com.
5 papers · 0 benchmarks
The Public Git Archive is a dataset of 182,014 top-bookmarked Git repositories from GitHub totalling 6 TB.
5 papers · 0 benchmarks
PyTorrent contains 218,814 Python package libraries from PyPI and Anaconda environment.
5 papers · 0 benchmarks
QAConv is a new question answering (QA) dataset that uses conversations as a knowledge source.
5 papers · 0 benchmarks
RARE (Randomized AMRs with Rewired Edges)
RARE consists of English AMR pairs with similarity scores that reflect the structural differences between them.
5 papers · 1 benchmark
RC-49 is a benchmark dataset for generating images conditional on a continuous scalar variable.
5 papers · 1 benchmark
This dataset arises from the READ project (Horizon 2020).
5 papers · 1 benchmark
RELiC is a large-scale dataset of 79k excerpts of literary scholarship, each containing a quotation from a primary source and the surrounding critical analysis.
5 papers · 0 benchmarks
The evaluation of object detection models is usually performed by optimizing a single metric, e.g.
5 papers · 1 benchmark
RITE (Retinal Images vessel Tree Extraction)
The RITE (Retinal Images vessel Tree Extraction) is a database that enables comparative studies on segmentation or classification of arteries and veins on retinal fundus images, which is established based on the public available DRIVE…
5 papers · 2 benchmarks
ROF (Real World Occluded Faces)
ROF is a dataset for occluded face recognition that contains faces with both upper face occlusion, due to sunglasses, and lower face occlusion, due to masks.
5 papers · 0 benchmarks
RRS Ranking Test (Restoration-200k for Response Selection with Ranking Test Set)
| | Train | Validation | Test | Ranking Test | | --------- | ----- | ---------- | ------- | ------------ | | size | 0.4M | 50K | 5K | 800 | | pos:neg | 1:1 | 1:9 | 1.2:8.8 | - | | avg turns | 5.0 | 5.0 | 5.0 | 5.0 | Ranking test set…
5 papers · 1 benchmark
RTL-Repo is a benchmark for evaluating LLMs' effectiveness in generating Verilog code autocompletions within large, complex codebases.
5 papers · 0 benchmarks
The RUN dataset is based on OpenStreetMap (OSM).
5 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
5 papers · 0 benchmarks
The Raider dataset collects fMRI recordings of 1000 voxels from the ventral temporal cortex, for 10 healthy adult participants passively watching the full-length movie “Raiders of the Lost Ark”.
5 papers · 0 benchmarks
This dataset has the following citation: M.
5 papers · 1 benchmark
RealMAN (A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization)
The Audio Signal and Information Processing Lab at Westlake University, in collaboration with AISHELL, has released the Real-recorded and annotated Microphone Array speech&Noise (RealMAN) dataset, which provides annotated multi-channel…
5 papers · 2 benchmarks
RefRef (RefRef: A Synthetic Dataset and Benchmark for Reconstructing Refractive and Reflective Objects)
RefRef is a synthetic dataset and benchmark designed for the task of reconstructing scenes with complex refractive and reflective objects.
5 papers · 1 benchmark
A dataset which contains over 200 apartments.
5 papers · 0 benchmarks
RiSAWOZ is a large-scale multi-domain Chinese Wizard-of-Oz dataset with Rich Semantic Annotations.
5 papers · 0 benchmarks
RoCoG-v2 (Robot Control Gestures) is a dataset intended to support the study of synthetic-to-real and ground-to-air video domain adaptation.
5 papers · 1 benchmark
The Russian Corpus of Linguistic Acceptability (RuCoLA) is built from the ground up under the well-established binary LA approach.
5 papers · 1 benchmark
RuCoS (Russian Reading Comprehension with Commonsense Reasoning)
Russian reading comprehension with Commonsense reasoning (RuCoS) is a large-scale reading comprehension dataset that requires commonsense reasoning.
5 papers · 1 benchmark
RuSentRel is a corpus of analytical articles translated into Russian texts in the domain of international politics obtained from foreign authoritative sources.
5 papers · 0 benchmarks
Ruddit is a dataset of English language Reddit comments that has fine-grained, real-valued scores for offensive language detection between -1 (maximally supportive) and 1 (maximally offensive).
5 papers · 0 benchmarks
S.MID (SeMantic InDustry)
SeMantic InDustry (S.MID) is a dataset designed to advance the field of LiDAR semantic segmentation, specifically for robotic applications and large-scale industrial scene.
5 papers · 1 benchmark
SAF (Short Answer Feedback Dataset)
This dataset can be found on HuggingFace: https://huggingface.co/datasets/Short-Answer-Feedback/safcommunicationnetworksenglish https://huggingface.co/datasets/Short-Answer-Feedback/safmicrojobgerman
5 papers · 0 benchmarks
SAFIM (Syntax-Aware Fill-In-the-Middle)
Syntax-Aware Fill-in-the-Middle (SAFIM) is a benchmark for evaluating Large Language Models (LLMs) on the code Fill-in-the-Middle (FIM) task.
5 papers · 1 benchmark
SDSD-indoor (Seeing Dynamic Scene in the Dark: High-Quality Video Dataset with Mechatronic Alignment)
The dataset collected by the paper Seeing Dynamic Scene in the Dark: High-Quality Video Dataset with Mechatronic Alignment, ICCV 2021
5 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.