Home › Datasets › task › regression
regression datasets
archive 2025-07-28
31 datasets carry the task tag "regression" (the task itself: regression), ordered by the archive's paper count. Page 1 of 1: 31 shown of 31. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
regression datasets 1–31 of 31
Satlas is a remote sensing dataset and benchmark that is large in both breadth, featuring all of the aforementioned applications and more, as well as scale, comprising 290M labels under 137 categories and 7 label modalities.
22 papers · 0 benchmarks
FLIP (Fitness Landscape Inference for Proteins)
FLIP includes several benchmark datasets that contain a variety of protein sequences, each with a real-valued label indicating its "fitness" (how well the protein performs some particular function).
11 papers · 0 benchmarks
SciRepEval is a comprehensive benchmark for training and evaluating scientific document representations.
9 papers · 0 benchmarks
Median house prices for California districts derived from the 1990 census.
6 papers · 2 benchmarks
Data variables and description.
3 papers · 0 benchmarks
ADORE (A benchmark dataset for machine learning in ecotoxicology)
ADORE is a benchmark dataset for machine learning for ecotixicology, covering acute aquatic toxicity in three relevant taxonomic groups (fish, crustaceans, and algae).
2 papers · 1 benchmark
Differential fluorescent staining is an effective tool widely adopted for the visualization, segmentation and quantification of cells and cellular substructures as a part of standard microscopic imaging protocols.
2 papers · 0 benchmarks
TTE-A&O (Travel Time Estimation: Abakan and Omsk)
The dataset includes two parts corresponding to the cities of Abakan (65524 nodes, 340012 edges) and Omsk (231688 nodes, 1149492 edges).
2 papers · 1 benchmark
bcTCGA (The Cancer Genome Atlas Program)
This data set comes from breast cancer tissue samples deposited to The Cancer Genome Atlas (TCGA) project.
2 papers · 0 benchmarks
From the official description: > The corpus contains 10-K reports from many US companies during years > 1996-2006, as well as measured volatility of stock returns for the > twelve-month periods preceding and following each report.
2 papers · 0 benchmarks
Unsustainable fishing practices worldwide pose a major threat to marine resources and ecosystems.
2 papers · 1 benchmark
In this dataset we added [Company Name, Car Model, Car Type, Fuel Type, Transmission, Engine (cc), Mileage, Kmsdriven, Buyers, Horsepower (kw), Year Price (Lakhs)]
1 paper · 1 benchmark
Concrete is the most important material in civil engineering.
1 paper · 1 benchmark
DRIFT (Domain-Adaptive Regression for Forest Monitoring)
The DRIFT dataset includes 25k image patches collected in five European countries sourced from aerial and nanosatellite image archives.
1 paper · 0 benchmarks
IOPS and Latency measurements of a real data storage system
1 paper · 0 benchmarks
FLIP includes several benchmark datasets that contain a variety of protein sequences, each with a real-valued label indicating its "fitness" (how well the protein performs some particular function).
1 paper · 0 benchmarks
Dataset Description This dataset contains rental property listings scraped from Tonaton.com, one of Ghana's leading online classifieds platforms.
1 paper · 0 benchmarks
Here the dataset described in Hitchhiking Rides Dataset: Two decades of crowd-sourced records on stochastic traveling(https://arxiv.org/abs/2506.21946) is published.
1 paper · 0 benchmarks
A representative event-based eye-tracking dataset, collected with two event cameras mounted on a glass frame.
1 paper · 2 benchmarks
We present two multi-modal datasets, one for Main Board IPOs, and the other for Small and Medium Enterprises (SME) IPOs.
1 paper · 0 benchmarks
LimeSoDa (Precision Liming Soil Datasets)
Precision Liming Soil Datasets (LimeSoDa) is a collection of 31 datasets from a field- and farm-scale soil mapping context.
1 paper · 0 benchmarks
MAX-60K (Masked Autoencoder for X-ray Fluorescence 60K Dataset)
The dataset for masked autoencoder for X-ray fluorescence (XRF) is a following development after the dataset (Chao et al., 2022).
1 paper · 0 benchmarks
This dataset contains demographic and personal health information for individuals, along with the corresponding medical insurance charges billed to them.
1 paper · 1 benchmark
Meta Omnium is a dataset-of-datasets spanning multiple vision tasks including recognition, keypoint localization, semantic segmentation and regression.
1 paper · 0 benchmarks
MiMIC (Multi-Modal Indian Earnings Calls Dataset)
Predicting stock market prices following corporate earnings calls remains a significant challenge for investors and researchers alike, requiring innovative approaches that can process diverse information sources.
1 paper · 0 benchmarks
We present a comprehensive dataset comprising a vast collection of raw mineral samples for the purpose of mineral recognition.
1 paper · 0 benchmarks
Dataset composed of two main parts 1.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The primary environmental health threat in the WHO European Region is air pollution, impacting the daily health and well-being of its citizens significantly.
1 paper · 0 benchmarks
Dataset Details This dataset is primarily created for the work Fast muon tracking with machine learning implemented in FPGA (Arxiv link) that contains ~3M simulated muon events with Geant4.
1 paper · 0 benchmarks
Enriched Voxceleb speakers' data of 1,715 celebrities with height gathered from Wikidata Other columns are taken from voxcelebenrichmentagegender
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.