Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 235 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 11233–11280 of 12,172
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
LLM-Based Vulnerability Classification in Police Narratives This repository contains datasets used in our research on applying large language models (LLMs) to identify indicators of vulnerability in police incident narratives.
1 paper · 0 benchmarks
VQA NLE synthetic dataset, made with LLaVA-1.5 using features from GQA dataset.
1 paper · 0 benchmarks
Dataset contains about 48K contracts which are open source on Etherscan.
1 paper · 0 benchmarks
Weather4Cast 2022 satellite images for weather prediction.
1 paper · 0 benchmarks
The dataset consists of 53,189 wikiHow articles across various categories of everyday tasks, 155,265 methods, and 772,294 steps with corresponding images.
1 paper · 1 benchmark
WikiEval Dataset for to do correlation analysis of difference metrics proposed in Ragas This dataset was generated from 50 pages from Wikipedia with edits post 2022.
1 paper · 0 benchmarks
Here I provided the datasets I used for this analysis.
1 paper · 0 benchmarks
WMT 2024 is a collection of datasets used in shared tasks of the Ninth Conference on Machine Translation.
1 paper · 0 benchmarks
A benchmark + dataset for evaluating multimodal models on business process management (BPM) tasks.
1 paper · 0 benchmarks
zbMATH Open contains over 4 million bibliographic entries with reviews or abstracts drawn from more than 3.000 journals and book series and more than 190.000 books.
1 paper · 0 benchmarks
This is a stance detection dataset in the Zulu language.
1 paper · 0 benchmarks
This is a collection of datasets related to Covid-19.
1 paper · 0 benchmarks
The dataset consists of tweets belonging to #MeToo movement on Twitter, labelled into different categories.
0 papers · 0 benchmarks
Demo dataset with 5 partial and complete 3D shapes of potato tubers.
0 papers · 0 benchmarks
Description: 4,458 People - 3D Facial Expressions Recognition Data.
0 papers · 0 benchmarks
This data collection consists of images acquired during chemoradiotherapy of 20 locally-advanced, non-small cell lung cancer patients.
0 papers · 0 benchmarks
The 5 deepnude apps that produce realistic and accurate results are mentioned below.
0 papers · 0 benchmarks
The world of adult content has been revolutionized by artificial intelligence, with AI porn generators pushing the boundaries of realism and creativity.
0 papers · 0 benchmarks
Description: 895 Fire Videos Data,the total duration of videos is 27 hours 6 minutes 48.58 seconds.
0 papers · 0 benchmarks
AASCE (Accurate Automated Spinal Curvature Estimation)
The purpose of this challenge is to investigate (semi-)automatic spinal curvature estimation algorithms.
0 papers · 0 benchmarks
ABODA (Abandoned Object Dataset)
ABandoned Objects DAtaset (ABODA) is a new public dataset for abandoned object detection.
0 papers · 0 benchmarks
ADFI (Anomaly Detection Datasets for Visual Inspection)
ADFI Dataset is an image dataset for anomaly detection methods with a focus on industrial inspection.
0 papers · 0 benchmarks
ADHA: “Adverbs Describing Human Actions” is the first benchmark for a new problem — recognizing human action adverbs (HAA).
0 papers · 0 benchmarks
AGGA (Academic Guidelines for Generative AIs) is a dataset of 80 academic guidelines for the usage of generative AIs and large language models in academia, selected systematically and collected from official university websites across six…
0 papers · 0 benchmarks
ALFI (Annotations for Label-Free Images)
ALFI (Annotations for Label-Free Images) is a dataset of images and annotations for label-free microscopy imaging.
0 papers · 0 benchmarks
ALT (Asian Language Treebank)
The ALT project aims to advance the state-of-the-art Asian natural language processing (NLP) techniques through the open collaboration for developing and using ALT.
0 papers · 0 benchmarks
This dataset is described in the ALTA 2022 Shared Task and associated CodaLab competition.
0 papers · 0 benchmarks
ALTA 2023 Shared Task (Discriminate between human-authored and synthetic text generated by Large Language Models (LLMs))
This dataset is described in the ALTA 2023 Shared Task and associated CodaLab competition.
0 papers · 0 benchmarks
AMUSE (Automotive Multi-Sensor Dataset)
The automotive multi-sensor (AMUSE) dataset consists of inertial and other complementary sensor data combined with monocular, omnidirectional, high frame rate visual data taken in real traffic scenes during multiple test drives.
0 papers · 0 benchmarks
Data Approximatively 2 hours of videos were captured from 7 viewpoints during a professional basketball game.
0 papers · 0 benchmarks
This is a new image-based handwritten historical digit dataset named ARDIS (Arkiv Digital Sweden).
0 papers · 0 benchmarks
ARF (Artificial Relationships in Fiction)
Artificial Relationships in Fiction Dataset Description Artificial Relationships in Fiction (ARF) is a synthetically annotated dataset for Relation Extraction (RE) in fiction, created from a curated selection of literary texts sourced from…
0 papers · 0 benchmarks
The first public dataset dedicated for Latin (French) and Arabic Scene Text Detection in Highway panels.
0 papers · 0 benchmarks
The ASL-Phono introduces a novel linguistics-based representation, which describes the signs in the ASLLVD dataset in terms of a set of attributes of the American Sign Language phonology.
0 papers · 0 benchmarks
The ASL-Skeleton3D introduces a representation based on mapping into the three-dimensional space the coordinates of the signers in the ASLLVD dataset.
0 papers · 0 benchmarks
This open-source dataset consists of 5.04 hours of transcribed English conversational speech beyond telephony, where 13 conversations were contained.
0 papers · 0 benchmarks
ASSIN (Avaliação de Similaridade Semântica e INferência textual) is a dataset with semantic similarity score and entailment annotations.
0 papers · 0 benchmarks
ASSIN2 (Avaliação de Similaridade Semântica e Inferência Textual) is the second edition of a workshop that evaluates Semantic Textual Similarity (STS) and Textual Entailment Recognition (RTE).
0 papers · 1 benchmark
The Aachen-Heerlen annotated steel microstructure dataset comprises 1,705 scanning electron microscopy (SEM) images of bainitic steel samples.
0 papers · 0 benchmarks
The Swedish ABSAbank-Imm 1.1 is an annotated corpus designed for aspect-based sentiment analysis related to immigration in Sweden.
0 papers · 0 benchmarks
This is the accompanying data and code for the publication [Markovitch & Krasnogor: Predicting Species Emergence in Simulated Complex Pre-Biotic Networks] containing the full set of 10,000 lognormal networks studied, their network…
0 papers · 0 benchmarks
Affective Text (Test Corpus of SemEval 2007) by Carlo Strapparava & Rada Mihalcea.
0 papers · 0 benchmarks
AiTLAS: Benchmark Arena is an open-source benchmark framework for evaluating state-of-the-art deep learning approaches for image classification in Earth Observation (EO).
0 papers · 0 benchmarks
AlexMI (Alex Motor Imagery dataset)
Alex Motor Imagery dataset.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.