Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 249 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 11905–11952 of 12,172
Consists of 36,785 images belonging to a diverse 92 classes.
0 papers · 0 benchmarks
Blockchain has empowered computer systems to be more secure using a distributed network.
0 papers · 0 benchmarks
As CryptoPunks pioneers the innovation of non-fungible tokens (NFTs) in AI and art, the valuation mechanics of NFTs has become a trending topic.
0 papers · 0 benchmarks
RepoQA is a benchmark that aims to exercise the long-context code understanding ability of Language Learning Models (LLMs)².
0 papers · 0 benchmarks
Files with responses by three different groups of participants to a questionnaire regarding preferences on humans vs artificial intelligence systems (or in support of humans) to make decisions.
0 papers · 0 benchmarks
This dataset contains a curated sample of 20,000 English-language user reviews sourced exclusively from Trustpilot.com.
0 papers · 0 benchmarks
Citation Request: See the articles for more detailed information on the data.
0 papers · 0 benchmarks
Dataset used for the challenge to apply computer vision techniques on art objects (paintings, sculptures, drawings etc) from the Rijksmuseum (in Amsterdam, the Netherlands).
0 papers · 0 benchmarks
Robbie Williams is a dataset of 65 songs by Robbie Williams.
0 papers · 0 benchmarks
Provides 8x 225-frame robotic surgical videos, captured at 2 Hz, where a trained team at Intuitive Surgical has manually labelled the different parts and types.
0 papers · 0 benchmarks
Based on Crawford’s work, we collect the most diverse and extensive image dataset of the reverse sides.
0 papers · 0 benchmarks
RuADReCT (The Russian Adverse Drug Reaction Corpus of Tweets)
Created as part of the Social Media Mining for Health Applications (#SMM4H '20) shared tasks, this dataset consists of 9515 tweets describing health issues.
0 papers · 0 benchmarks
RuFa (Ruqaa-Farsi) dataset contains images of text written in one of two Arabic fonts: Ruqaa and Nastaliq (Farsi).
0 papers · 0 benchmarks
This dataset contains annotations of semantic frames and intra-frame syntax for 1500 Russian sentences.
0 papers · 0 benchmarks
A runway dataset, designing features suitable for capturing outfit appearance, collecting human judgments of outfit similarity, and learning similarity functions on the features to mimic those judgments.
0 papers · 0 benchmarks
SAR Patches (SAR Multiple Times, Polarization And Orbital Pass)
Sentinel-1 SAR samples from a set of manually chosen points across the world from different times polarization and orbital pass.
0 papers · 0 benchmarks
SBWCE (Spanish Billion Word Corpus and Embeddings)
This resource consists of an unannotated corpus of the Spanish language of nearly 1.5 billion words, compiled from different corpora and resources from the web; and a set of word vectors (or embeddings), created from this corpus using the…
0 papers · 0 benchmarks
SC-artext (https://huggingface.co/datasets/SSS/SC-artext)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
0 papers · 0 benchmarks
SEN (Sentiment analysis of Entities in News headlines)
SEN is a novel publicly available human-labelled dataset for training and testing machine learning algorithms for the problem of entity level sentiment analysis of political news headlines.
0 papers · 0 benchmarks
Dataset Card for SENTINEL: Mitigating Object Hallucinations via Sentence-Level Early Intervention For the details of this dataset, please refer to the documentation of the GitHub repo.
0 papers · 0 benchmarks
SESYD "Systems Evaluation SYnthetic Documents" is a database of synthetical documents with groundtruth.
0 papers · 0 benchmarks
Facial landmark detection is a cornerstone in many facial analysis tasks such as face recognition, drowsiness detection, and facial expression recognition.
0 papers · 0 benchmarks
SICS-155 (Phase Recognition in Small Incision Cataract Surgery Videos)
Cataract is the leading cause of blindness worldwide, most affecting life in low- and middle-income countries (LMICs).
0 papers · 0 benchmarks
This dataset was taken from the SIGARRA information system at the University of Porto (UP).
0 papers · 0 benchmarks
SILD (Survey Item Linking Dataset)
This dataset contains a collection of texts from publications from a broad range of social science domains (e.g., economics, politics, psychology, etc.).
0 papers · 0 benchmarks
SIT (Symbol Interpretation Task)
This task is composed of five different subtasks that require interpreting statements referring to structures of a simple world.
0 papers · 0 benchmarks
The SKEMPI database contains data on the changes in thermodynamic parameters and kinetic rate constants upon mutation, for protein-protein interactions for which a structure of the complex has been solved and is available in the Protein…
0 papers · 0 benchmarks
SMAC+ defensive infantry scenario with sequential episodic buffer
0 papers · 0 benchmarks
SMCOVID19-CT (Contact Tracing Data (from Italian SM-COVID-19 App))
We present a real data analysis of a CT experiment that was conducted in Italy for 8 months and involved more than 100,000 CT app users.
0 papers · 0 benchmarks
SMDG (Standardized Multi-Channel Dataset for Glaucoma)
Standardized Multi-Channel Dataset for Glaucoma (SMDG-19) is a collection and standardization of 19 public datasets, comprised of full-fundus glaucoma images, associated image metadata like, optic disc segmentation, optic cup segmentation,…
0 papers · 0 benchmarks
Seizures and seizure-like rhythmic and periodic brain activity known as “ictal-interictal-injury continuum” (IIIC) patterns are frequently detected during brain monitoring with electroencephalography (EEG) in patients with epilepsy or…
0 papers · 0 benchmarks
STAR is a novel benchmark for Situated Reasoning, which provides 60K challenging situated questions in four types of tasks, 140K situated hypergraphs, symbolic situation descriptions and logic-grounded diagnosis for real-world video…
0 papers · 0 benchmarks
Grounding Scientific Entity References in STEM Scholarly Content to Authoritative Encyclopedic and Lexicographic Sources The STEM ECR v1.0 dataset has been developed to provide a benchmark for the evaluation of scientific entity…
0 papers · 0 benchmarks
STVD-FC is the largest public dataset on the political content analysis and fact-checking tasks.
0 papers · 0 benchmarks
SUC (Stockholm-Umeå Corpus)
The Stockholm-Umeå Corpus (SUC) is a collection of Swedish texts from the 1990s, consisting of one million words in total.
0 papers · 0 benchmarks
SUIM-E (SGUIE-Net: Semantic attention guided underwater image enhancement with multi-scale perception)
Underwater Image Enhancement Dataset
0 papers · 0 benchmarks
Saarbruecken Voice Database contains voice and EGG recordings of patients diagnosed with voice disorder, as well as healthy persons.
0 papers · 0 benchmarks
The Safety Prompts dataset is a valuable resource for evaluating and enhancing the safety of large language models (LLMs) in the Chinese language.
0 papers · 0 benchmarks
Sakha-TB (400+400 CXR images for TB diagnosis)
Sakha-TB is a de-identified image dataset of frontal chest X-rays (CXR), collected through collaboration with several medical institutions in the Republic of Sakha (Yakutia, Russia).
0 papers · 0 benchmarks
The shared task ScienceIE at SemEval 2017 deals with automatic extraction of keyphrases from Computer Science, Material Sciences and Physics publications, as well as extracting types of keyphrases and relations between keyphrases.
0 papers · 0 benchmarks
The Second HAREM was an evaluation exercise in Portuguese Named Entity Recognition.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.