Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 231 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 11041–11088 of 12,172
This repository contains BasqueParl, a bilingual corpus for political discourse analysis.
1 paper · 0 benchmarks
In this repository, we provide the set-up files and output files of 5 behavioral observation data entry applications.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Bespoke-Stratos-17k We replicated and improved the Berkeley Sky-T1 data pipeline using SFT distillation data from DeepSeek-R1 to create Bespoke-Stratos-17k -- a reasoning dataset of questions, reasoning traces, and answers.
1 paper · 0 benchmarks
Downloadable zip file containing raw data (simulated and real) as well as model fits / saved parameters.
1 paper · 0 benchmarks
A Collection of Multilingual Parallel Datasets for 5 Indonesian Local Languages
1 paper · 0 benchmarks
bigscience/P3 (bigscience/P3, split='ai2_arc_ARC_Challenge_pick_the_most_correct_option')
This datasets consists of challenging reasoning questions in multiple choice format.
1 paper · 0 benchmarks
Microarray gene expression data on 57 bladder samples from 5 batches.
1 paper · 0 benchmarks
blbooks (The British Library Books)
This dataset consists of books digitised by the British Library in partnership with Microsoft.
1 paper · 0 benchmarks
sloptmagmajsons.tar.gz contains the summary of the Magma benchmark (Section 5.4) as JSON files, which was generated by exp2json.py.
1 paper · 0 benchmarks
The cCOVID-News dataset is a publicly available Chinese text retrieval dataset created from COVID-19 news articles.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The controller area network (CAN) bus has emerged as the de facto standard for in-vehicle networks (IVNs) around the globe.
1 paper · 0 benchmarks
We aim to improve the bAbI benchmark as a means of developing intelligent dialogue agents.
1 paper · 1 benchmark
We aim to improve the bAbI benchmark as a means of developing intelligent dialogue agents.
1 paper · 1 benchmark
sloptfuzzbenchandbanditplotdata.tar.gz contains all plotdata of fuzzer instances that were run in the FuzzBench benchmark (Section 5.3) and Bandit Algorithm Comparison (Section 4.3).
1 paper · 0 benchmarks
A fully synthetic dataset of drones generated using structured domain randomization.
1 paper · 0 benchmarks
cryoPPP (CryoPPP: A Large Expert-Curated Cryo-EM Image Dataset for Machine Learning Protein Particle Picking)
The CryoPPP dataset consists of 34 ground truth data and metadata for 335 EMPIAR IDs.
1 paper · 0 benchmarks
dHCP (developing Human Connectome Project)
The dHCP dataset contains neonatal MRI.
1 paper · 0 benchmarks
Download free fonts in DaFont style from our extensive collection.
1 paper · 0 benchmarks
Included in this content: 0045.perovksitedata.csv - main dataset used in this article.
1 paper · 0 benchmarks
data_qe (Federal Reserve Quantitative Easing Data)
This file contains the data and code for the publication "The Federal Reserve's Response to the Global Financial Crisis and Its Long-Term Impact: An Interrupted Time-Series Natural Experimental Analysis" by A.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Debatepedia is a debate platform that lists arguments to a topic on one page, including subtitles, structuring the arguments into different aspects.
1 paper · 1 benchmark
Release decompile-ghidra-100k, a subset of 100k training samples (25k per optimization level).
1 paper · 0 benchmarks
deepMTJ (Muscle-Tendon Junction Tracking in Ultrasound Images)
deepMTJ: Muscle-Tendon Junction Tracking in Ultrasound Images ------------------------------------------------------------- deepMTJ is a machine learning approach for automatically tracking of muscle-tendon junctions (MTJ) in ultrasound…
1 paper · 1 benchmark
This dataset contains manipulated images and real images.
1 paper · 0 benchmarks
This dataset addresses the lack of public botnet datasets, especially for the IoT.
1 paper · 0 benchmarks
Dataset for our paper Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky which includes 5000 enterprise tools and the corresponding dialogues generated using DiaFORGE UTC data engine.
1 paper · 0 benchmarks
Diderot’s Encyclopédie is a reference work from XVIIIth century in Europe that aimed at collecting the knowledge of its era.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Perturbed version of the data used in Van Dijcke, Gunsilius, and Wright (2024).
1 paper · 0 benchmarks
Perturbed version of the data used in Van Dijcke, Gunsilius, and Wright (2024).
1 paper · 0 benchmarks
This is the list of all doges of the Venetian Republic, as well as their wives, if there's a record that they existed.
1 paper · 0 benchmarks
The database is comprised of two datasets, the 4-Class eTRIMS Dataset with 4 annotated object classes and the 8-Class eTRIMS Dataset with 8 annotated object classes.
1 paper · 0 benchmarks
eVED (Extended Vehicle Energy Dataset)
Extended Vehicle Energy Dataset (eVED) is an extended version of the Vehicle Energy Dataset (VED), which is a large-scale dataset for vehicle energy consumption analysis.
1 paper · 0 benchmarks
ec-darkpattern is a dataset for dark pattern detection and prepared its baseline detection performance with state-of-the-art machine learning methods.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
f5C Dataset (5-Formylcytidine Modifications on mRNA Dataset)
This is a dataset for predicting 5-Formylcytidine Modifications on mRNA.
1 paper · 1 benchmark
fGn series used in the article to develop the simulations.
1 paper · 0 benchmarks
fake (Real / Fake Job Posting Prediction)
[Real or Fake] : Fake Job Description Prediction This dataset contains 18K job descriptions out of which about 800 are fake.
1 paper · 1 benchmark
A collection of natural language prompt-completion pairs pertaining to multiple-choice Q&A on benchmark tasks based on US census products.
1 paper · 0 benchmarks
IMO-level geometry problem with complete natural language description, geometric shapes, formal language annotations, and theorem sequences annotations.
1 paper · 0 benchmarks
fruit-SALAD is a synthetic image dataset with 10,000 generated images of fruit depictions.
1 paper · 0 benchmarks
gComm is a step towards developing a robust platform to foster research in grounded language acquisition in a more challenging and realistic setting.
1 paper · 0 benchmarks
gENder-IT is an English-Italian challenge set focusing on the resolution of natural gender phenomena by providing word-level gender tags on the English source side and multiple gender alternative translations, where needed, on the Italian…
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.