Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 233 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 11137–11184 of 12,172
This dataset is composed of paired videos of people dancing 3 different music styles: Ballet, Michael Jackson and Salsa.
1 paper · 0 benchmarks
This is the Big-Bench version of our language-based movie recommendation dataset https://github.com/google/BIG-bench/tree/main/bigbench/benchmarktasks/movierecommendation GPT-2 has a 48.8% accuracy, chance is 25%.
1 paper · 1 benchmark
This dataset links all the entries describing named entities of Petit Larousse illustré, a French dictionary published in 1905, to wikidata identifiers.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
lilGym is a benchmark for language-conditioned reinforcement learning in visual environment based on 2,661 highly-compositional human-written natural language statements grounded in an interactive visual environment.
1 paper · 0 benchmarks
This repository contains the code and data to reproduce all results in "Climate uncertainty impacts on social cost of carbon and optimal mitigation pathways", Smith et al.
1 paper · 0 benchmarks
To construct our multilingual dataset - mBBC - we gathered news articles from various BBC news websites in 43 different languages.
1 paper · 0 benchmarks
mDRT (Multilingual Diagnostic Rhyme Test)
We present a multilingual test set for conducting speech intelligibility tests in the form of diagnostic rhyme tests.
1 paper · 0 benchmarks
Multilingual Open Knowledge Base Completion benchmark in 6 languages: English, Hindi, Telugu, Spanish, Portuguese, and Chinese.
1 paper · 0 benchmarks
mdCATH (mdCATH: A Large-Scale MD Dataset for Data-Driven Computational Biophysics)
This dataset comprises all-atom systems for 5,398 CATH domains, modeled with a state-of-the-art classical force field, and simulated in five replicates each at five temperatures from 320 K to 450 K.
1 paper · 0 benchmarks
medisim is a collection of new large-scale medical term similarity datasets based on SNOMED-CT.
1 paper · 0 benchmarks
This dataset is used to examine the influence of image memorability on social media virality.
1 paper · 0 benchmarks
Item-wise accuracies in six benchmarks from Open LLM Leaderboard 1 scraped from huggingface.co and used for metabench analyses and construction.
1 paper · 0 benchmarks
micro-emotion dataset (Supplementary annotated Goemotions dataset (micro-emotion labels with energy level intensity values (0-10))
The EQN framework is a micro-emotion annotation and detection system that realizes the automatic micro-emotion annotation of text with energy level scores for the first time.
1 paper · 0 benchmarks
🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages.
1 paper · 0 benchmarks
🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages.
1 paper · 0 benchmarks
🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages.
1 paper · 0 benchmarks
Military Vehicles for Hierarchical Multi-label Classification (HMC)
1 paper · 0 benchmarks
We introduce misinfo-general, a benchmark dataset for evaluating misinformation models’ ability to perform out-of-distribution generalisation.
1 paper · 0 benchmarks
A modification on the ShEMO dataset with help of an Automatic Speech Recognition (ASR) system.
1 paper · 0 benchmarks
the dataset is a monkey doo doo dataset
1 paper · 0 benchmarks
If you want to known more about this dataset and new method, please read our paper link.
1 paper · 0 benchmarks
To encourage reproducible research, a labeled MultiRAW dataset containing>7k RAW images acquired using multiple camera sensors is made publicly accessible for RAW-domain processing.
1 paper · 0 benchmarks
mwBTFreddy dataset is a resource developed to support flash flood damage assessment in urban Malawi, specifically focusing on the impacts of Cyclone Freddy in 2023.
1 paper · 0 benchmarks
naab: A ready-to-use plug-and-play corpus for Farsi The biggest cleaned and ready-to-use open-source textual corpus in Farsi.
1 paper · 0 benchmarks
needadvice is a dataset for advice classification extracted from Reddit.
1 paper · 0 benchmarks
see detailed descriptions in readme.md
1 paper · 0 benchmarks
This dataset contains 17,090 audio clips of length 30 seconds sampled from archives collected from 6 Guinean radio stations.
1 paper · 0 benchmarks
This dataset contains 10,083 recorded utterances in French, Maninka, Pular and Susu from 49 speakers (16 female and 33 male) ranging from 5 to 76 years old on a variety of devices.
1 paper · 0 benchmarks
Noise event recordings with static scenes.
1 paper · 0 benchmarks
A new SIS benchmark designed to assess generation performance under noisy conditions, simulating human error that can occur during real-world applications.
1 paper · 1 benchmark
A new SIS benchmark designed to assess generation performance under noisy conditions, simulating human error that can occur during real-world applications.
1 paper · 1 benchmark
A new SIS benchmark designed to assess generation performance under noisy conditions, simulating human error that can occur during real-world applications.
1 paper · 1 benchmark
A cross-city UDA benchmark built upon nuScenes.
1 paper · 0 benchmarks
[Dataset on HF] [Project Page] [Subjective LeaderBoard] [Objective LeaderBoard] CriticBench is a novel benchmark designed to comprehensively and reliably evaluate the critique abilities of Large Language Models (LLMs).
1 paper · 0 benchmarks
A collection of datasets and benchmarks for large-scale Performance Modeling with LLMs.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The dataset describes 150 patients with the following demographic characteristics : sex, age, HOMA-IR , systolic and diastolic blood pressure , and LDL-Cholesterol .
1 paper · 0 benchmarks
pd4ml (Physics Data for Machine Learning)
pd4ml is a collection of datasets from fundamental physics research -- including particle physics, astroparticle physics, and hadron- and nuclear physics -- for supervised machine learning studies.
1 paper · 0 benchmarks
pglib-opf (Power Grid Lib - Optimal Power Flow)
This benchmark library is curated and maintained by the IEEE PES Task Force on Benchmarks for Validation of Emerging Power System Algorithms and is designed to evaluate a well established version of the the AC Optimal Power Flow problem.
1 paper · 0 benchmarks
The pic2kal benchmark for calorie prediction contains 308,000 images from over 70,000 recipes including photographs, ingredients and instructions, matched with nutritional information.
1 paper · 0 benchmarks
In this dataset we teleoperated UR5 arm to collect manipulation data for picking up a screwdriver in a cluttered tabletop environment.
1 paper · 0 benchmarks
pmuBAGE (the Benchmarking Assortment of Generated PMU Events) is a dataset that consists of almost 1000 instances of labeled event data to encourage benchmark evaluations on phasor measurement unit (PMU) data analytics.
1 paper · 0 benchmarks
Political stance in Danish.
1 paper · 0 benchmarks
This dataset tests the capabilities of language models to correctly capture the meaning of words denoting probabilities (WEP, also called verbal probabilities), e.g.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.