Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 177 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 8449–8496 of 12,172
A new large-scale dataset that consists of 409 fine-grained categories and 31,881 images with accurate 3D pose annotation.
1 paper · 0 benchmarks
FineCops-Ref is a dataset for Compositional Referring Expression Comprehension (REC) that rigorously evaluates Vision-Language Models (VLMs) on compositional reasoning and their ability to identify inconsistencies between images and text.
1 paper · 0 benchmarks
This dataset is part of the Finetune-RAG project, which aims to tackle hallucination in retrieval-augmented LLMs.
1 paper · 0 benchmarks
Synthetic training set: This set is constructed in the following two steps and will be used for estimation/training purposes.
1 paper · 0 benchmarks
This dataset comprises 1-minute fingertip video recordings collected from 150 anemic patients, ranging from 6 months to 32 years of age, with hemoglobin levels between 4.3 gm/dL and 12.4 gm/dL.
1 paper · 0 benchmarks
Fire Drill Anti-Pattern Dataset is a collection of nine real-world software projects for detection of the fire drill anti-pattern with ground truth, issue-tracking data, source code density, models and code.
1 paper · 0 benchmarks
This dataset is collected by DataCluster Labs, India.
1 paper · 1 benchmark
HAREM, an initiative by Linguateca, boasts a Golden Collection—a meticulously curated repository of annotated Portuguese texts.
1 paper · 0 benchmarks
The researchers collected a dataset of 3,500 images of Tilapia fish in a small bowl containing three fish per image.
1 paper · 0 benchmarks
The researchers collected 3,500 images of Tilapia fish, with each image containing three fish in a small bowl.
1 paper · 0 benchmarks
We introduce FixEval , a dataset for competitive programming bug fixing along with a comprehensive test suite and show the necessity of execution based evaluation compared to suboptimal match based evaluation metrics like BLEU, CodeBLEU,…
1 paper · 0 benchmarks
Spectrum data on CH4/Air flame emission with 200ms exposure and 2s exposure
1 paper · 0 benchmarks
This dataset contains: (1) Slforge Generated Simulink Models : Synthetic Simulink Models (2) Source of Real World Simulink Models The .txt file is a combined text file that contains all the real world Simulink models based on SLGPT's…
1 paper · 0 benchmarks
The Flipkart Products Review Dataset is a collection of reviews provided by customers who purchased products from Flipkart, one of India’s leading e-commerce platforms.
1 paper · 0 benchmarks
FloCo (Flow chart Image to Code)
the FloCo dataset that contains 11,884 flowchart images and their corresponding Python codes.
1 paper · 1 benchmark
The Florentine dataset is a dataset of facial gestures which contains facial clips from 160 subjects (both male and female), where gestures were artificially generated according to a specific request, or genuinely given due to a shown…
1 paper · 0 benchmarks
This dataset contains several instances of the Offline Nanosatellite Task Scheduling (ONTS) problem, based on the parameters of the FloripaSat-1 mission.
1 paper · 0 benchmarks
MDA231 human breast carcinoma cells infected with a pMSCV vector including the GFP sequence, embedded in a collagen matrix Dr.
1 paper · 1 benchmark
Simulated nuclei of HL60 cells stained with Hoescht Dr.
1 paper · 1 benchmark
The Foggy KITTI dataset extends the KITTI dataset to include challenging weather conditions, aiming to support research in real-world applications such as autonomous driving.
1 paper · 0 benchmarks
This dataset is a result of a study that was created to assess drivers behaviors when following a lead vehicle.
1 paper · 0 benchmarks
FFR Dataset is an ongoing project to collect, clean and store corpora of Fon and French sentences for machine translation from Fon-French.
1 paper · 0 benchmarks
Fongbe Data collected by Fréjus A.
1 paper · 1 benchmark
This data set encompasses 104 images and transcriptions of digital images of original charters from the Cistercian abbey Fontenay in Burgundy (France), dating mainly from the 12th c.
1 paper · 0 benchmarks
FooDI-ML (Food Drinks and groceries Images Multi Lingual)
Food Drinks and groceries Images Multi Lingual (FooDI-ML) is a dataset that contains over 1.5M unique images and over 9.5M store names, product names descriptions, and collection sections gathered from the Glovo application.
1 paper · 2 benchmarks
About Dataset The file contains 24K unique figure obtained from various Google resources Meticulously curated figure ensuring diversity and representativeness Provides a solid foundation for developing robust and precise figure allocation…
1 paper · 0 benchmarks
The Food Recall Incidents dataset consists of 7,546 short texts (from 5 to 360 characters each), which are the titles of food recall announcements (therefore referred to as title), crawled from 24 public food safety authority websites by…
1 paper · 0 benchmarks
Food.com Recipes and Interactions consists of 270K recipes and 1.4M user-recipe interactions (reviews) scraped from Food.com, covering a period of 18 years (January 2000 to December 2018).
1 paper · 0 benchmarks
FoodSG-233 (Localized Singaporean Food Image Dataset)
The FoodSG-233 dataset contains 209,861 images, covering 13 food groups and 233 food categories.
1 paper · 0 benchmarks
ForPKG (https://github.com/luozhongze/ForPKG)
A policy knowledge graph can provide decision support for tasks such as project compliance, policy analysis, and intelligent question answering, and can also serve as an external knowledge base to assist the reasoning process of related…
1 paper · 0 benchmarks
Here is the forbidden question dataset (based on two previous works), it contains 160 questions from 160 violated categories.
1 paper · 0 benchmarks
Here is the forbidden question dataset (based on two previous works), it contains 160 questions from 160 violated categories.
1 paper · 0 benchmarks
This is a simulated dataset for force prediction
1 paper · 0 benchmarks
Repository for the question sets and resolution sets described produced by ForecastBench, a forecasting benchmark for LLMs.
1 paper · 0 benchmarks
This dataset contains news headlines relevant to key forex pairs: AUDUSD, EURCHF, EURUSD, GBPUSD, and USDJPY.
1 paper · 0 benchmarks
FormAI is a novel AI-generated dataset comprising 112,000 compilable and independent C programs.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This is an extension to the studyforrest dataset – a versatile resource for studying the behavior of the human brain in situations of real-life complexity (http://studyforrest.org).
1 paper · 0 benchmarks
Fraud_Case_Verdicts (The "Crime Facts" of "Offenses of Fraudulence" in Judicial Yuan Verdicts Dataset)
The "Crime Facts" of "Offenses of Fraudulence" in Judicial Yuan Verdicts Dataset This data set is based on the judgments of "Offenses of Fraudulence" cases published by the Judicial Yuan.
1 paper · 0 benchmarks
The 'Me 163' was a Second World War fighter airplane and a result of the German air force secret developments.
1 paper · 0 benchmarks
The Fraunhofer Portugal AICOS EDoF Dataset was produced within the TAMI project and is composed of images of microscopic fields of view (FOV) of Liquid-based Cervical Cytology (LBC) samples.
1 paper · 0 benchmarks
FreCDo is a corpus for French dialect identification comprising 413,522 French text samples collected from public news websites in Belgium, Canada, France and Switzerland.
1 paper · 0 benchmarks
Free Law Project is a leading nonprofit organization that aims to make the legal ecosystem more equitable and competitive through technology, data, and advocacy.
1 paper · 1 benchmark
FreeMan is the first large-scale multi-view human motion dataset under real scenarios.
1 paper · 0 benchmarks
The Freesound One-Shot Percussive Sounds dataset contains 10254 one-shot (single event) percussive sounds from Freesound.org and the corresponding timbral analysis.
1 paper · 0 benchmarks
The Freiburg Street Crossing dataset consists of data collected from three different street crossings in Freiburg, Germany; ; two of which were traffic light regulated intersections and one a zebra crossing without traffic lights.
1 paper · 0 benchmarks
Freiburg Terrains consist of three parts: 3.7 hours of audio recordings of the microphone pointed at the robot wheels.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.