Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 63 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 2977–3024 of 12,172
WiderPerson contains a total of 13,382 images with 399,786 annotations, i.e., 29.87 annotations per image, which means this dataset contains dense pedestrians with various kinds of occlusions.
11 papers · 1 benchmark
These data are the results of a chemical analysis of wines grown in the same region in Italy but derived from three different cultivars.
11 papers · 6 benchmarks
5,519 query-based summaries, each associated with an average of 6 input documents selected from an index of 355M documents from Common Crawl.
11 papers · 0 benchmarks
comma 2k19 is a dataset of over 33 hours of commute in California's 280 highway.
11 papers · 0 benchmarks
iMiGUE is a dataset for emotional artificial intelligence research: identity-free video dataset for Micro-Gesture Understanding and Emotion analysis (iMiGUE).
11 papers · 1 benchmark
methods2test is a supervised dataset consisting of Test Cases and their corresponding Focal Methods from a set of Java software repositories.
11 papers · 0 benchmarks
Robust detection and tracking of objects is crucial for the deployment of autonomous vehicle technology.
11 papers · 2 benchmarks
nvBench is a large-scale NL2VIS (natural languagge to visualisations) benchmark, containing 25,750 (NL, VIS) pairs from 750 tables over 105 domains, synthesized from (NL, SQL) benchmarks to support cross-domain NLPVIS (Natural Language…
11 papers · 0 benchmarks
2-PM Vessel is an open-source volumetric brain vasculature dataset obtained with two-photon microscopy at Focused Ultrasound Lab, at Sunnybrook Research Institute (affiliated with University of Toronto by Dr.
10 papers · 0 benchmarks
The AI City Challenge, hosted at CVPR 2024, focuses on harnessing AI to enhance operational efficiency in physical settings such as retail and warehouse environments, and Intelligent Traffic Systems (ITS).
10 papers · 1 benchmark
Novel benchmark which features aspects of natural scenes, e.g.
10 papers · 1 benchmark
A benchmark for matching and registration of partial point clouds with time-varying geometry.
10 papers · 1 benchmark
ABCD (Action-Based Conversations Dataset)
10 papers · 1 benchmark
ABIDE (Autism Brain Imaging Data Exchange)
Autism spectrum disorder (ASD) is characterized by qualitative impairment in social reciprocity, and by repetitive, restricted, and stereotyped behaviors/interests.
10 papers · 0 benchmarks
We release expert-made scribble annotations for the medical ACDC dataset [1].
10 papers · 1 benchmark
AKB-48 is a large-scale Articulated object Knowledge Base which consists of 2,037 real-world 3D articulated object models of 48 categories.
10 papers · 0 benchmarks
ATD-12K is a large-scale animation triplet dataset, which comprises 12,000 triplets(train10k,test2k) by manually inspect and the test2k with rich annotations, including levels of difficulty, the Regions of Interest (RoIs) on movements, and…
10 papers · 1 benchmark
Contains densely labeled speech activity in YouTube videos, with the goal of creating a shared, available dataset for this task.
10 papers · 1 benchmark
AnnoMI: A Dataset of Expert-Annotated Counselling Dialogues Dataset Introduction Research on natural language processing approaches to analysing counselling dialogues has seen substantial development in recent years, but access to this…
10 papers · 0 benchmarks
Bank Account Fraud (BAF) is a large-scale, realistic suite of tabular datasets.
10 papers · 12 benchmarks
The goal of the "BCI Competition" is to validate signal processing and classification methods for Brain-Computer Interfaces (BCIs).
10 papers · 0 benchmarks
BiRD (Bigram Relatedness Dataset)
Bigram Relatedness Dataset (BiRD) is a large, fine-grained, bigram relatedness dataset, using a comparative annotation technique called Best Worst Scaling.
10 papers · 0 benchmarks
BigDatasetGAN is a dataset for pixel-wise ImageNet segmentation.
10 papers · 0 benchmarks
A binarized version of MNIST.
10 papers · 1 benchmark
BindingDB is a public, web-accessible database of measured binding affinities, focusing chiefly on the interactions of protein considered to be drug-targets with small, drug-like molecules.
10 papers · 2 benchmarks
CAPE (Clothed Auto Person Encoding)
The CAPE dataset is a 3D dynamic dataset of clothed humans, featuring: - 3D mesh registrations of accurate scans of clothed people in motion, captured at 60 FPS; - Consistent SMPL mesh topology, all frames in correspondence; - Precise,…
10 papers · 1 benchmark
An expert-annotated word similarity dataset which provides a highly reliable, yet challenging, benchmark for rare word representation techniques.
10 papers · 0 benchmarks
CC-Stories (or STORIES) is a dataset for common sense reasoning and language modeling.
10 papers · 0 benchmarks
CDDB (Continual Deepfake Detection Benchmark)
Abstract: There have been emerging a number of benchmarks and techniques for the detection of deepfakes.
10 papers · 0 benchmarks
CHALET (Cornell House Agent Learning Environment)
CHALET is a 3D house simulator with support for navigation and manipulation.
10 papers · 0 benchmarks
CJRC (Chinese judicial reading comprehension)
The Chinese judicial reading comprehension (CJRC) dataset contains approximately 10K documents and almost 50K questions with answers.
10 papers · 0 benchmarks
CLEVR-Dialog is a large diagnostic dataset for studying multi-round reasoning in visual dialog.
10 papers · 0 benchmarks
CMeEE (Chinese Medical Named Entity Recognition Dataset)
Chinese Medical Named Entity Recognition, a dataset first released in CHIP20204, is used for CMeEE task.
10 papers · 1 benchmark
A large-scale curated dataset of over 152 million tweets, growing daily, related to COVID-19 chatter generated from January 1st to April 4th at the time of writing.
10 papers · 0 benchmarks
Curates a large pelvic CT dataset pooled from multiple sources and different manufacturers, including 1, 184 CT volumes and over 320, 000 slices with different resolutions and a variety of the above-mentioned appearance variations.
10 papers · 0 benchmarks
CVCS (Cross-View Cross-Scene Multi-View Crowd Counting Dataset)
CVCS is a synthetic multi-view people dataset, containing 31 scenes, where 23 are for training and the rest 8 for testing.
10 papers · 1 benchmark
CalMS21 (Caltech Mouse Social Interactions)
The Caltech Mouse Social Interactions (CalMS21) dataset is a multi-agent dataset from behavioral neuroscience.
10 papers · 0 benchmarks
Chart2Text is a dataset that was crawled from 23,382 freely accessible pages from statista.com in early March of 2020, yielding a total of 8,305 charts, and associated summaries.
10 papers · 0 benchmarks
ChatHaruhi (ChatHaruhi: Reviving Anime Character in Reality via Large Language Model)
ChatHaruhi is a dataset covering 32 Chinese / English TV / anime characters with over 54k simulated dialogues.
10 papers · 0 benchmarks
ChestX-Det is a chest X-Ray dataset with instance-level annotations (boxes and masks).
10 papers · 0 benchmarks
Consist of 23,533 statements extracted from all U.S.
10 papers · 0 benchmarks
Java-Small, Java-Med, Java-Large
10 papers · 0 benchmarks
CommitPackFT is a 2GB filtered version of CommitPack to contain only high-quality commit messages that resemble natural language instructions.
10 papers · 0 benchmarks
The CommitmentBank is a corpus of 1,200 naturally occurring discourses whose final sentence contains a clause-embedding predicate under an entailment canceling operator (question, modal, negation, antecedent of conditional).
10 papers · 1 benchmark
CriticBench is a comprehensive benchmark designed to assess the abilities of Large Language Models (LLMs) to critique and rectify their reasoning across various tasks.
10 papers · 0 benchmarks
The CropAndWeed dataset is focused on the fine-grained identification of 74 relevant crop and weed species with a strong emphasis on data variability.
10 papers · 0 benchmarks
CubiCasa5K is a large-scale floorplan image dataset containing 5000 samples annotated into over 80 floorplan object categories.
10 papers · 0 benchmarks
To enrich the diversity, we also collect 92 images which are suitable for saliency detection from DAVIS [27], a densely annotated high-resolution video segmentation dataset.
10 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.