Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 149 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 7105–7152 of 12,172
WeatherKITTI is currently the most realistic all-weather simulated enhancement of the KITTI dataset.
2 papers · 0 benchmarks
We developed Web-Bench as a benchmark for evaluating the performance of LLMs on real-world web projects.
2 papers · 1 benchmark
This paper is a condensed report on the second year of the Touché shared task on argument retrieval held at CLEF 2021.
2 papers · 0 benchmarks
Weibo-COV is a large-scale COVID-19 social media dataset from Weibo, covering more than 30 million posts from 1 November 2019 to 30 April 2020.
2 papers · 0 benchmarks
WetLinks (WetLinks: a Large-Scale Longitudinal Starlink Dataset with Contiguous Weather Data)
WetLinks: a Large-Scale Longitudinal Starlink Dataset with Contiguous Weather Data.
2 papers · 0 benchmarks
WiFall (Wireless Sensing Dataset for Fall Detection, Action Recognition and People ID Identification with ESP32-S3)
WiFall dataset contains data related to fall detection, action recognition and people id identification in a meeting room scenario.
2 papers · 1 benchmark
WikiCaps is a large-scale multilingual but non-parallel data set for multimodal machine translation and retrieval.
2 papers · 0 benchmarks
A newly developed public dataset and the task of multiple property extraction.
2 papers · 0 benchmarks
WikiSRS is a novel dataset of similarity and relatedness judgments of paired Wikipedia entities (people, places, and organizations), as assigned by Amazon Mechanical Turk workers.
2 papers · 0 benchmarks
WikiTableSet is a large publicly available image-based table recognition dataset in three languages built from Wikipedia.
2 papers · 0 benchmarks
Wikidata-14M is a recommender system dataset for recommending items to Wikidata editors.
2 papers · 0 benchmarks
The Wikidata-Disamb dataset is intended to allow a clean and scalable evaluation of NED with Wikidata entries, and to be used as a reference in future research.
2 papers · 0 benchmarks
Contains Wikipedia pages about popular mathematics topics and edges describe the links from one page to another.
2 papers · 0 benchmarks
WildDESED (Wild Domestic Environment Sound Event Detection)
WildDESED is an extension of the original DESED dataset, created to reflect various domestic scenarios by incorporating complex and unpredictable background noises.
2 papers · 1 benchmark
WildPPG (WildPPG: A Real-World PPG Dataset of Long Continuous Recordings)
a dataset of multi-modal signals from wearable devices at four sites on the body.
2 papers · 1 benchmark
This is a dataset for the task of PE-type malware in the Windows operating system.
2 papers · 0 benchmarks
Wyze Rule Recommendation Dataset.
2 papers · 0 benchmarks
X-WikiRE is a new, large-scale multilingual relation extraction dataset in which relation extraction is framed as a problem of reading comprehension to allow for generalization to unseen relations.
2 papers · 0 benchmarks
It consists of an extensive collection of a high quality cross-lingual fact-to-text dataset in 11 languages: Assamese (as), Bengali (bn), Gujarati (gu), Hindi (hi), Kannada (kn), Malayalam (ml), Marathi (mr), Oriya (or), Punjabi (pa),…
2 papers · 1 benchmark
XImageNet (XIMAGENET-12: An Explainable AI Benchmark Dataset for Model Robustness Evaluation)
we introduce XIMAGENET-12, an explainable benchmark dataset with over 200K images and 15,600 manual semantic annotations.
2 papers · 0 benchmarks
XL-R2R (Cross-lingual Room-to-Room)
The XL-R2R dataset is built upon the R2R dataset and extends it with Chinese instructions.
2 papers · 0 benchmarks
Given a question and passage in an Indic language, generate a short answer span from the passage as the answer.
2 papers · 0 benchmarks
Given a question in an Indic language and a passage in English, generate a short answer span.
2 papers · 0 benchmarks
YASO is a crowd-sourced TSA evaluation dataset, collected using a new annotation scheme for labeling targets and their sentiments.
2 papers · 0 benchmarks
We present YTSeg, a topically and structurally diverse benchmark for the text segmentation task based on YouTube transcriptions.
2 papers · 1 benchmark
This dataset contains 94 movie summary videos from various YouTube channels.
2 papers · 0 benchmarks
The YouTube8M-MusicTextClips dataset consists of over 4k high-quality human text descriptions of music found in video clips from the YouTube8M dataset.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
This repository contains the Zurich Transit Bus (ZTBus) dataset, which consists of data recorded during driving missions of electric city buses in Zurich, Switzerland.
2 papers · 0 benchmarks
This is a set of datasets containing three versions of data: - V0: the original sampled addresses with no augmentation.
2 papers · 0 benchmarks
arXivEdits an annotated corpus of 751 full papers from arXiv with gold sentence alignment across their multiple versions of revision, as well as fine-grained span-level edits and their underlying intentions for 1,000 sentence pairs.
2 papers · 0 benchmarks
bcTCGA (The Cancer Genome Atlas Program)
This data set comes from breast cancer tissue samples deposited to The Cancer Genome Atlas (TCGA) project.
2 papers · 0 benchmarks
The bipedal skills benchmark is a suite of reinforcement learning environments implemented for the MuJoCo physics simulator.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
ccHarmony is a color checker (cc) based image harmonization dataset.
2 papers · 0 benchmarks
dacl10k (dacl10k: Dataset for Semantic Bridge Damage Segmentation)
dacl10k stands for damage classification 10k images and is a multi-label semantic segmentation dataset for 19 classes (13 damages and 6 objects) present on bridges.
2 papers · 0 benchmarks
This dataset comprises 1344 expert annotated images of muscle-tendon junctions recorded with 3 ultrasound imaging systems (Aixplorer V6, Esaote MyLab60, Telemed ArtUs), on 2 muscles (Lateral Gastrocnemius, Medial Gastrocnemius), and 2…
2 papers · 0 benchmarks
This dataset contains pre and post destruction images and also segmentation labels for test images.
2 papers · 0 benchmarks
dichasus-cf0x (CSI Dataset dichasus-cf0x: Distributed Antenna Setup in Industrial Environment, Day 1)
Dataset containing channel state information (CSI) alongside ground truth data (position tags, timestamps) of a massive MIMO-OFDM system measured with the DICHASUS channel sounder.
2 papers · 0 benchmarks
From the official description: > The corpus contains 10-K reports from many US companies during years > 1996-2006, as well as measured volatility of stock returns for the > twelve-month periods preceding and following each report.
2 papers · 0 benchmarks
ePillID is a benchmark for developing and evaluating computer vision models for pill identification.
2 papers · 0 benchmarks
Overview The edeniss2020 dataset is a time series dataset.
2 papers · 0 benchmarks
- A set of java bugs - With executable test cases
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.