Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 176 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 8401–8448 of 12,172
FROG is a 2D LiDAR dataset with annotations for people detectors.
1 paper · 0 benchmarks
FSC-P2 (Fearless Steps Challenge Phase2)
The Fearless Steps Initiative by UTDallas-CRSS led to the digitization, recovery, and diarization of 19,000 hours of original analog audio data, as well as the development of algorithms to extract meaningful information from this…
1 paper · 0 benchmarks
A synthetic sound mixture specification dataset for the Target Sound Extraction (TSE) task.
1 paper · 2 benchmarks
FSOCO is a collaborative dataset for vision-based cone detection systems in Formula Student Driverless competitions.
1 paper · 0 benchmarks
FT Speech is a speech corpus created from the recorded meetings of the Danish Parliament, otherwise known as the Folketing (FT).
1 paper · 0 benchmarks
FTR-18 is a multilingual rumour dataset on football transfer news.
1 paper · 0 benchmarks
FUSU dataset covers 5 whole urban areas, 847 km^2 located in the north and south of China, with 17 land use and land cover (LULC) classes and over 170K images and 30 billion pixels of annotations, supporting segmentation, change detection…
1 paper · 0 benchmarks
A set of 248 search queries annotated with the correct diagnosis.
1 paper · 0 benchmarks
Faces Through Time (FTT) features 26,247 images of notable people from the 19th to 21st centuries, with roughly 1,900 images per decade on average.
1 paper · 0 benchmarks
Facial Skeletal angles (Facial Skeletal Angles (Glabella and Maxilla Angle and Length and Width of Piriformis))
Facial Skeletal Angles (Glabella and Maxilla Angle and Length and Width of Piriformis)
1 paper · 0 benchmarks
A Racial Fairness Benchmark Dataset for Face Forgery Detection.
1 paper · 0 benchmarks
The FairTranslate Dataset includes 2,418 sentence pairs, each centered around an occupation, designed to assess gender expression and translation in English-French contexts.
1 paper · 0 benchmarks
Expertly-curated benchmark dataset for fake news detection in Filipino.
1 paper · 0 benchmarks
Fallout New Vegas Dialog is a multilingual sentiment annotated dialog dataset from Fallout New Vegas.
1 paper · 0 benchmarks
The "Famous Keyword Twitter Replies Dataset" is a comprehensive collection of Twitter data that focuses on popular keywords and their associated replies.
1 paper · 0 benchmarks
A dataset of fashion images for social events To collect Fashion4Events, a dataset of garment images paired with social event labels, we exploited two different sources of data: the DeepFashion2 dataset and the USED dataset.
1 paper · 0 benchmarks
The FashionFail dataset comprises 2,495 high-resolution images (2400x2400 pixels) of products found on e-commerce websites.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Copyright (C) 2021 Ante Qu .
1 paper · 0 benchmarks
Structure of code/data folders and how to use them fastzip-code Contains codebase to generate results in fastzip-results folder Individual notebooks contain comments on their functionality/how to use them FastZIP-Resample.ipynb (optional)…
1 paper · 0 benchmarks
Dataset with raw outputs of experiments connected to the GitHub repository: https://github.com/yurilavinas/MOEADr/tree/ECJ The size of the data is big, make sure to have enough space.
1 paper · 0 benchmarks
The FathomNet2023 competition dataset is a subset of the broader FathomNet marine image repository.
1 paper · 0 benchmarks
The FeatherV1 dataset is a dataset for fine-grained visual classification.
1 paper · 0 benchmarks
FedNLP (FOMC Docs and Speeches)
We collect the various forms of Federal Reserve communications.
1 paper · 0 benchmarks
FedTADBench is a federated time series anomaly detection benchmark.
1 paper · 0 benchmarks
DOI: https://doi.org/10.7910/DVN/O4CRXK The most comprehensive standardised data on Malaysian federal and state elections from 1955 to the present.
1 paper · 0 benchmarks
This folder contains 5-arcmin resolution maps for the application rate of each fertilizer (N,P2O5,K2O) on cropland and each of the 13 major crop groups resulting from our study.
1 paper · 0 benchmarks
FewDR is a dataset for Few-shot dense retrieval (DR).
1 paper · 0 benchmarks
Introduction The FewGLUE64labeled dataset is a new version of FewGLUE dataset.
1 paper · 0 benchmarks
FiVE (A Fine-grained Video Editing Benchmark)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This dataset is a "part II" extension of the "Engineered cardiac microbundle time-lapse microscopy image dataset" and contains 808 experimental time-lapse image sequences of beating hiPSC-based cardiac microbundles using FibroTUG platforms…
1 paper · 0 benchmarks
File S1 (Carboxylase properties table)
-Tab 1 (Carboxylase table): This expanded table contains additional information for carboxylase classes and splits them into individual examples.
1 paper · 0 benchmarks
-Tab 1 Rubisco forms: This excel sheet contains one row for every rubisco form considered in this review (some forms like IAq and IAc from [31] are not considered separately because they are only phylogenetically separated in the small…
1 paper · 0 benchmarks
This .csv file contains all of the sequences used in the phylogenetic analysis (see above).
1 paper · 0 benchmarks
File S4 (tree file of all rubiscos at 65% dereplication)
This tree was generated as indicated above in the methods.
1 paper · 0 benchmarks
File S5 (tree file of Form IV rubiscos (RLPs) at 70% dereplication)
This tree was generated as indicated above in the methods.
1 paper · 0 benchmarks
File S6 (tree file of Form I rubiscos at 85% dereplication)
This tree was generated as indicated above in the text for file S3 using model LG+F+G.
1 paper · 0 benchmarks
File S7 (tree file of Form I rubiscos at 85% dereplication LG+R8)
This tree was generated as indicated above in the methods.
1 paper · 0 benchmarks
Filipino CrowS-Pairs and Filipino WinoQueer assess sexist and homophobic biases in language models handling Filipino.
1 paper · 0 benchmarks
A large film style dataset
1 paper · 0 benchmarks
FilmStills is a dataset of stills taken from a variety of films and TV shows, each concatenated with a color-compressed (with a factor of 2.667) version of itself.
1 paper · 0 benchmarks
48 multitrack jazz recordings with many annotations.
1 paper · 2 benchmarks
FinChat (Finnish Chat Conversations on Everyday Topics)
Finnish chat conversation corpus and includes unscripted conversations on seven topics from people of different ages.
1 paper · 0 benchmarks
Enhancing Financial Market Predictions: Causality-Driven Feature Selection This paper introduces FinSen dataset that revolutionizes financial market analysis by integrating economic and financial news articles from 197 countries with stock…
1 paper · 1 benchmark
FinDKG: The Global Financial Dynamic Knowledge Graph Dataset FinDKG is an open-source dataset focused on creating a temporally-resolved Financial Dynamic Knowledge Graph.
1 paper · 0 benchmarks
Financial Language Understanding Evaluation is an open-source comprehensive suite of benchmarks for the financial domain.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.