Home › Datasets › modality › Tabular
Tabular datasets
archive 2025-07-28
267 datasets carry the modality tag "Tabular", ordered by the archive's paper count. Page 2 of 6: 48 shown of 267. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Tabular datasets 49–96 of 267
eICU-CRD (eICU Collaborative Research Database)
The eICU Collaborative Research Database is a large multi-center critical care database made available by Philips Healthcare in partnership with the MIT Laboratory for Computational Physiology.
4 papers · 2 benchmarks
The eSports Sensors dataset contains sensor data collected from 10 players in 22 matches in League of Legends.
4 papers · 2 benchmarks
The CTU Relational Learning Repository offers relational database datasets to the machine learning community.
3 papers · 0 benchmarks
The dataset contains historical technical data of Dhaka Stock Exchange (DSE).
3 papers · 0 benchmarks
A new spatio-temporal benchmark dataset (Hurricane), is suited for forecasting during extreme events and anomalies.
3 papers · 1 benchmark
The GitTables-SemTab dataset is a subset of the GitTables dataset and was created to be used during the SemTab challenge.
3 papers · 2 benchmarks
HELOC (Home Equity Line of Credit)
HELOC The HELOC dataset from FICO.
3 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 0 benchmarks
A MIDI dataset of 500 4-part chorales generated by the KSChorus algorithm, annotated with results from hundreds of listening test participants, with 500 further unannotated chorales.
3 papers · 0 benchmarks
The Kvasir-VQA dataset is an extended dataset derived from the HyperKvasir and Kvasir-Instrument datasets, augmented with question-and-answer annotations.
3 papers · 0 benchmarks
Open Dataset: Mobility Scenario FIMU An open, multidimensional (6 categorical attributes), and synthetic dataset of faked virtual humans generated by an optimization approach applied to a real-life call-detail-records-based anonymized…
3 papers · 0 benchmarks
The Papers with Code Leaderboards dataset is a collection of over 5,000 results capturing performance of machine learning models.
3 papers · 1 benchmark
Dataset Overview This dataset contains individual-level data from a randomized controlled trial (RCT) conducted in northern Uganda, along with associated satellite imagery.
3 papers · 0 benchmarks
Retweet MTPP (Marked Temporal Point Processes on Retweet data)
This dataset contains time-stamped user retweet event sequences.
3 papers · 1 benchmark
SupplyGraph (SupplyGraph: A Benchmark Dataset for Supply Chain Planning using Graph Neural Networks)
Graph Neural Networks (GNNs) have gained traction across different domains such as transportation, bio-informatics, language processing, and computer vision.
3 papers · 0 benchmarks
This dataset contains aircraft trajectories in an untowered terminal airspace collected over 8 months surrounding the Pittsburgh-Butler Regional Airport [ICAO:KBTP], a single runway GA airport, 10 miles North of the city of Pittsburgh,…
3 papers · 1 benchmark
Travel (Tour & Travels Customer Churn Prediction)
A Tour & Travels Company Wants To Predict Whether A Customer Will Churn Or Not Based On Indicators Given Below.
3 papers · 1 benchmark
In our benchmark WHYSHIFT, we explore distribution shifts on 5 real-world tabular datasets from the economic and traffic sectors with natural spatiotemporal distribution shifts.We only pick 7 typical settings out of 22 settings and select…
3 papers · 0 benchmarks
The collected dataset consists of multivariate time series (MTS) data belonging to several ATMs banking along with the annotations that the operators did when they performed a maintenance task on any of the machines.
2 papers · 0 benchmarks
Measurement data related to the publication „Active TLS Stack Fingerprinting: Characterizing TLS Server Deployments at Scale“.
2 papers · 0 benchmarks
The dataset contains historical financial transactions, including time, category and cost fields.
2 papers · 1 benchmark
Amazon MTPP (Marked Temporal Point Processes on Amazon data)
The dataset includes time-stamped user product reviews behavior from January, 2008 to October, 2018.
2 papers · 1 benchmark
BASEPROD (The Bardenas Semi-Desert Planetary Rover Dataset)
BASEPROD provides comprehensive rover sensor data collected over a 1.7 km traverse, accompanied by high-resolution 2D and 3D drone maps of the terrain.
2 papers · 0 benchmarks
The Berlin V2X dataset offers high-resolution GPS-located wireless measurements across diverse urban environments in the city of Berlin for both cellular and sidelink radio access technologies, acquired with up to 4 cars over 3 days.
2 papers · 0 benchmarks
Several datasets are fostering innovation in higher-level functions for everyone, everywhere.
2 papers · 0 benchmarks
A coronavirus dataset with 98 countries constructed from different reliable sources, where each row represents a country, and the columns represent geographic, climate, healthcare, economic, and demographic factors that may contribute to…
2 papers · 0 benchmarks
This experiment was performed in order to empirically measure the energy use of small, electric Unmanned Aerial Vehicles (UAVs).
2 papers · 1 benchmark
This repository contains the database of the FEM simulation of axially impacted various configurations of the square crash boxes.
2 papers · 0 benchmarks
FINDSum (Financial Report Document Summarization)
FINDSum is a large-scale dataset for long text and multi-table summarization.
2 papers · 0 benchmarks
FinBench is a benchmark for evaluating the performance of machine learning models with both tabular data inputs and profile text inputs.
2 papers · 0 benchmarks
GIRT-Data (GitHub Issue Report Template Dataset)
GIRT-Data is the first and largest dataset of issue report templates (IRTs) in both YAML and Markdown format.
2 papers · 0 benchmarks
Hotel (Hospitality > Tourism > Hotel Demand/Sales)
The dataset contains the hotel demand and revenue of 8 major tourist destinations in the US (e.g., Los Angeles, Orlando ...).
2 papers · 0 benchmarks
Timely and effective response to humanitarian crises requires quick and accurate analysis of large amounts of text data, a process that can highly benefit from expert-assisted NLP systems trained on validated and annotated data in the…
2 papers · 0 benchmarks
IHDS (Indian Human Developement Survey)
IHDS is a nationally representative, multi-topic panel survey of 41,554 households in 1503 villages and 971 urban neighborhoods across India.
2 papers · 0 benchmarks
Replication Data for: Integrating Earth Observation Data into Causal Inference: Challenges and Opportunities Details: YandWmat.csv contains individual-level observational data.
2 papers · 0 benchmarks
The Insider Threat Test Dataset is a collection of synthetic insider threat test datasets that provide both background and malicious actor synthetic data.
2 papers · 1 benchmark
This dataset presents a set of large-scale ridesharing Dial-a-Ride Problem (DARP) instances.
2 papers · 0 benchmarks
This dataset contains the extraction made in 2022 of all the 622 datasets that existed then at the UCI Machine Learning Repository.
2 papers · 0 benchmarks
The original dataset was provided by Orange telecom in France, which contains anonymized and aggregated human mobility data.
2 papers · 0 benchmarks
News Interactions on Globo.com (News Portal User Interactions by Globo.com - A large dataset for news recommendations offline evaluation and analytics)
Context This large dataset with users interactions logs (page views) from a news portal was kindly provided by [Globo.com][1], the most popular news portal in Brazil, for reproducibility of the experiments with CHAMELEON - a…
2 papers · 0 benchmarks
Dataset Card for The Cancer Genome Atlas (TCGA) Multimodal Dataset The Cancer Genome Atlas (TCGA) Multimodal Dataset is a comprehensive collection of clinical data, pathology reports, molecular, and slide images for cancer patients.
2 papers · 0 benchmarks
A dataset consisting of recipient 46 users and, 26180 tweets.
2 papers · 0 benchmarks
Articles originating from subreddits with explicitly stated ideologies are categorized into three groups: 72,488 articles in the Liberal class, 79,573 articles in the Conservative class, and 225,083 articles in the Restricted class.
2 papers · 1 benchmark
Transaction fee mechanism (TFM) is an essential component of a blockchain protocol.
2 papers · 0 benchmarks
SNDZoo (The Softwarised Network Data Zoo)
The softwarised network data zoo (SNDZoo) is an open collection of software networking data sets aiming to streamline and ease machine learning research in the software networking domain.
2 papers · 0 benchmarks
The resources for this dataset can be found at https://www.openml.org/d/182 Author: Ashwin Srinivasan, Department of Statistics and Data Modeling, University of Strathclyde Source: UCI - 1993 Please cite: UCI The database consists of the…
2 papers · 0 benchmarks
This resource is designed to allow for research into Natural Language Generation.
2 papers · 0 benchmarks
The dataset has two years of user awards on a question-answering website: each user received a sequence of badges and there are 22 different kinds of badges in total.
2 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.