Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 152 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 7249–7296 of 12,172
AAAC (Artificial Argument Analysis Corpus)
DeepA2 is a modular framework for deep argument analysis.
1 paper · 0 benchmarks
AADB2021 (Cation-coordinated conformers of 20 proteinogenic amino acids with different protonation states)
We present a data set from a first-principles study of amino-methylated and acetylated (capped) dipeptides of the 20 proteinogenic amino acids – including alternative possible side chain protonation states and their interactions with…
1 paper · 0 benchmarks
AADB2021Ontology (Ontology representation for a data set of cation-coordinated conformers of 20 proteinogenic amino acids)
This onotology is populated with the data from AADB2021 (https://dx.doi.org/10.17172/NOMAD/2021.02.10-1).
1 paper · 0 benchmarks
AAVE/SAE Paired Dataset contains 2019 intent-equivalent AAVE/SAE pairs.
1 paper · 0 benchmarks
ABCD Study (Adolescent Brain Cognitive Development)
The ABCD Study is a prospective longitudinal study starting at the ages of 9-10 and following participants for 10 years.
1 paper · 0 benchmarks
ABUZZ (Citizen-based mosquito monitoring system)
As part of our policy to openly share all data from this project, we have included a downloadable package comprising all acoustic data collected over the course of this work.
1 paper · 0 benchmarks
AC-Bench: A Benchmark for Actual Causality Reasoning Dataset Description AC-Bench is designed to evaluate the actual causality (AC) reasoning capabilities of large language models (LLMs).
1 paper · 0 benchmarks
A benchmark environment based on the datasets "Adult" and "Names", which allows researchers to test how well their language model can abide by pre-defined access rights rules.
1 paper · 0 benchmarks
ACCORD CSQA is an extension of the popular CommonsenseQA (CSQA) dataset using ACCORD, a scalable framework for disentangling the commonsense grounding and reasoning abilities of large language models (LLMs) through controlled, multi-hop…
1 paper · 0 benchmarks
ACCT Data Repository (ACCT is a fast and accessible automatic cell counting tool using machine learning for 2D image segmentation)
This dataset is a collection of fluorescent images from mice in order to test an automatic cell counting tool that we developed.
1 paper · 0 benchmarks
ACFR Orchard Fruit Dataset is an agricultural dataset containing images and annotations for different fruits, collected at different farms across Australia.
1 paper · 0 benchmarks
This repository provides full-text and metadata to the ACL anthology collection (80k articles/posters as of September 2022) also including .pdf files and grobid extractions of the pdfs.
1 paper · 0 benchmarks
ACLUE (Ancient Chinese Language Understanding Evaluation)
The Ancient Chinese Language Understanding Evaluation (ACLUE) is an evaluation benchmark focused on ancient Chinese language comprehension.
1 paper · 0 benchmarks
Dataset Card for the ACR Appropriateness Criteria Corpus This dataset contains chunked guidelines and narratives from the ACR Appropriateness Criteria, an set of societal guidelines from the American College of Radiology (ACR) to help…
1 paper · 0 benchmarks
This data set contains 50 low resolution (640 x 360) short videos containing a variety real life activities.
1 paper · 0 benchmarks
This is a comprehensive dataset of human arm motion during Activities of Daily Living (ADL).
1 paper · 0 benchmarks
Three target attributes like AD123, ABETA12, and AV45AB12, representing various stages ofAlzheimer’s disease and captured through DTI analysis for white matter integrity.
1 paper · 0 benchmarks
AE (answer correctness dataset)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
AESI (Athens Emotional States Inventory)
The development of ecologically valid procedures for collecting reliable and unbiased emotional data towards computer interfaces with social and affective intelligence targeting patients with mental disorders.
1 paper · 0 benchmarks
AG-VPReID (AG-VPReID: A Challenging Large-Scale Benchmark for Aerial-Ground Video-based Person Re-Identification)
The largest aerial-ground video dataset with 6,632 identities, 32,321 tracklets, and 9.6 million frames Multiple platforms combining aerial drones (at various altitudes: 15m, 30m, 80m, 120m), CCTV cameras, and wearable cameras Rich…
1 paper · 0 benchmarks
AGB-DE is a legal NLP corpus for the automated detection of potentially void clauses in German standard form consumer contracts.
1 paper · 1 benchmark
AI Generated Content (AIGC) refers to any form of content, such as text, images, audio, or video, that is created with the help of artificial intelligence technology.
1 paper · 0 benchmarks
Consists of 7.5k sentences with gapping (as well as 15k relevant negative sentences) and comprises data from various genres: news, fiction, social media and technical texts.
1 paper · 0 benchmarks
Replication Material This document contains the necessary materials and instructions to replicate the findings presented in our paper.
1 paper · 0 benchmarks
Despite the availability of vast amounts of data, legal data is often unstructured, making it difficult even for law practitioners to ingest and comprehend the same.
1 paper · 0 benchmarks
The AI-GA (Artificial Intelligence Generated Abstracts) dataset is a collection of abstracts and titles, with half of the abstracts being AI-generated and the other half being original.
1 paper · 0 benchmarks
This project contains instructions and codes to reconstruct a dataset for the development and evaluation of forensic tools for detecting machine-generated text in social media.
1 paper · 0 benchmarks
We introduce a new AI-ready computational pathology dataset containing restained and co-registered digitized images from eight head-and-neck squamous cell carcinoma patients.
1 paper · 0 benchmarks
AIDERV2 (Aerial Image Dataset for Emergency Response Applications (version 2))
The dataset contains aerial images containing three commonly occurring natural disasters earthquake/collapsed buildings, flood, wildfire/fire, and a normal class; do not reflect any disaster.
1 paper · 1 benchmark
The AIDS Antiviral Screen dataset is a dataset of screens checking tens of thousands of compounds for evidence of anti-HIV activity.
1 paper · 0 benchmarks
AIROGS (Rotterdam EyePACS AIROGS)
The Rotterdam EyePACS AIROGS dataset (in full, so including train and test) contains 113,893 color fundus images from 60,357 subjects and approximately 500 different sites with a heterogeneous ethnicity.
1 paper · 0 benchmarks
AISECKG (AISecKG: Knowledge Graph Dataset for Cybersecurity Education)
Cybersecurity education is exceptionally challenging as it involves learning the complex attacks; tools and developing critical problem-solving skills to defend the systems.
1 paper · 0 benchmarks
In AISIA-VN-Review-S and AISIA-VN-Review-F datasets, we first collect 450K customer reviewing comments from various e–commerce websites.
1 paper · 0 benchmarks
Synthetic log data suitable for evaluation of intrusion detection systems, federated learning, and alert aggregation.
1 paper · 0 benchmarks
Video samples recorded in the field using the Azure Kinect DK.
1 paper · 0 benchmarks
ALLO (Anomaly Localization in Lunar Orbit)
ALLO is an anomaly detection and localization dataset for space stations in lunar orbit.
1 paper · 0 benchmarks
ALPHA (ALPHA: AnomaLous Physiological Health Assessment Using Large Language Models)
This study concentrates on evaluating the efficacy of Large Language Models (LLMs) in healthcare, with a specific focus on their application in personal anomalous health monitoring.
1 paper · 0 benchmarks
we collected a new real-world dataset, called ALPIXVSR, using a ALPIX-Eiger event camera1 .
1 paper · 0 benchmarks
AMFDS (Arabic Multi-Fonts Dataset)
Arabic Multi Fonts Dataset A multi-word multi-font Arabic word-image dataset.
1 paper · 0 benchmarks
The AML Robot Cutting Dataset consists of approximately 1500 seconds of real data collected on Kinova Jaco 2 robot retrofitted with a custom end-effector fixture and dremel performing cutting tasks on wood specimens for 5 materials and 5…
1 paper · 0 benchmarks
AMR3.0 (Abstract Meaning Representation (AMR) Annotation Release 3.0)
Abstract Meaning Representation (AMR) Annotation Release 3.0 was developed by the Linguistic Data Consortium (LDC), SDL/Language Weaver, Inc., the University of Colorado's Computational Language and Educational Research group and the…
1 paper · 1 benchmark
AMT Objects is a large dataset of object centric videos suitable for training and benchmarking models for generating 3D models of objects from a small number of photos of the objects.
1 paper · 0 benchmarks
ANCHOLIK-NER (ANCHOLIK-NER: A Benchmark Dataset for Bangla Regional Named Entity Recognition)
We developed ANCHOLIK-NER, a Bangla Regional Named Entity Recognition dataset focusing on the Sylhet, Chittagong, and Barishal dialects.
1 paper · 0 benchmarks
ANFC (A Comprehensive Dataset and Automated Pipeline for Nailfold Capillary Analysis)
Nailfold capillaroscopy stands as a traditional and classical method for health condition assessment.
1 paper · 0 benchmarks
It comprises synthetic mesh sequences from Deformation Transfer for Triangle Meshes.
1 paper · 0 benchmarks
ANTILLES (ANTILLES: An Open French Linguistically Enriched Part-of-Speech Corpus)
ANTILLES is a part-of-speech tagging corpus based on UDFrench-GSD which was originally created in 2015 and is based on the universal dependency treebank v2.0.
1 paper · 1 benchmark
ANUBIS (Skeleton-Based Action Recognition Dataset)
ANUBIS is a large-scale human skeleton dataset containing 80 actions.
1 paper · 0 benchmarks
AODRaw (Adverse condition Object Detection with RAW images)
We introduce the AODRaw dataset, which offers 7,785 high-resolution real RAW images with 135,601 annotated instances spanning 62 categories, capturing a broad range of indoor and outdoor scenes under 9 distinct light and weather conditions.
1 paper · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.