Home › Datasets › modality › Tabular
Tabular datasets
archive 2025-07-28
267 datasets carry the modality tag "Tabular", ordered by the archive's paper count. Page 5 of 6: 48 shown of 267. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Tabular datasets 193–240 of 267
This dataset are about Nafion 112 membrane standard tests and MEA activation tests of PEM fuel cell in various operation condition.
1 paper · 0 benchmarks
The data set includes information about 120+ elections (configuration settings and descriptive statistics), projects and 125k+ anonymized voters and their budget preferences.
1 paper · 0 benchmarks
This repository contains a dataset and machine learning algorithms to detect poisoned water from clean water via using equivalent Smartphone embedded Wi-Fi CSI data.
1 paper · 0 benchmarks
Engagement with the government of Taiwan as part of the vTaiwan participatory process which led to the successful regulation of Uber in Taiwan.
1 paper · 0 benchmarks
PreRAID (Prescreening Rheumatoid Arthritis Information Database (PreRAID))
PreRAID is a structured dataset designed to evaluate the diagnostic capabilities of Large Language Models (LLMs) in Rheumatoid Arthritis (RA) diagnosis.
1 paper · 0 benchmarks
PsOCR (Pashto OCR Dataset)
PsOCR is a large-scale synthetic dataset for Optical Character Recognition in low-resource Pashto language.
1 paper · 0 benchmarks
We create a new dataset from GitTables, a data lake of 1.7M tables extracted from CSV files on GitHub.
1 paper · 0 benchmarks
We conducted a large crowdsourcing study of click patterns in an interactive segmentation scenario and collected 475K real-user clicks.
1 paper · 0 benchmarks
RFSD (Russian Financial Statements Database)
The Russian Financial Statements Database (RFSD) The Russian Financial Statements Database (RFSD) is an open, harmonized collection of annual unconsolidated financial statements of the universe of Russian firms.
1 paper · 0 benchmarks
The data used in - "Radio Galaxy Zoo EMU: Towards a Semantic Radio Galaxy Morphology Taxonomy" (Bowles et al.
1 paper · 0 benchmarks
The following files contains the simulation inputs and outputs for conducting the multi-objetive optimization of thermal comfort and dyalight with the Response Surface Methodology.
1 paper · 0 benchmarks
Teaching assistants (TAs) are heavily used in computer science courses as a way to handle high enrollment and still being able to offer students individual tutoring and detailed assessments.
1 paper · 0 benchmarks
This dataset was acquired in a retrospective study from a cohort of pediatric patients admitted with abdominal pain to Children’s Hospital St.
1 paper · 0 benchmarks
Bitcoin is a peer-to-peer electronic payment system that popularized rapidly in recent years.
1 paper · 0 benchmarks
ata Set Name: Rice Dataset (Commeo and Osmancik) Abstract: A total of 3810 rice grain's images were taken for the two species (Cammeo and Osmancik), processed and feature inferences were made.
1 paper · 0 benchmarks
Fact-based Text Editing dataset based on RotoWire dataset
1 paper · 1 benchmark
Our SRSD (Feynman) datasets are designed to discuss the performance of Symbolic Regression for Scientific Discovery.
1 paper · 0 benchmarks
Our SRSD (Feynman) datasets are designed to discuss the performance of Symbolic Regression for Scientific Discovery.
1 paper · 0 benchmarks
Our SRSD (Feynman) datasets are designed to discuss the performance of Symbolic Regression for Scientific Discovery.
1 paper · 0 benchmarks
Dataset Card for SemTabNet This dataset accompanies the following paper: Title: Statements: Universal Information Extraction from Tables with Large Language Models for ESG KPIs Authors: Lokesh Mishra, Sohayl Dhibi, Yusik Kim, Cesar…
1 paper · 1 benchmark
This data comprises processed weather, soil, yield, and cultivation area for corn yield prediction in Sub-Sahara Africa, with emphasis on Nigeria.
1 paper · 0 benchmarks
The SNS data (Valente et al., 2013) is a four-wave survey conducted in Los Angeles county, the United States, which features a sample of 1,795 high-school students.
1 paper · 0 benchmarks
Songdo Traffic (Songdo Traffic: High Accuracy Georeferenced Vehicle Trajectories from a Large-Scale Study in a Smart City)
The Songdo Traffic dataset delivers precisely georeferenced vehicle trajectories captured through high-altitude bird's-eye view (BeV) drone footage over Songdo International Business District, South Korea.
1 paper · 0 benchmarks
Songdo Vision (Songdo Vision: Vehicle Annotations from High-Altitude BeV Drone Imagery in a Smart City)
The Songdo Vision dataset provides high-resolution (4K, 3840×2160 pixels) RGB images annotated with categorized axis-aligned bounding boxes (BBs) for vehicle detection from a high-altitude bird’s-eye view (BeV) perspective.
1 paper · 1 benchmark
Classifying Email as Spam or Non-Spam.
1 paper · 0 benchmarks
This dataset consists of EEG (Electroencephalogram) recordings collected from students at our college during an educational experiment.
1 paper · 0 benchmarks
The file contains an annotated list of papers that are included in the literature survey.
1 paper · 0 benchmarks
Survey answers (Answers to surveys in both papers, as well as processed answers)
Please see paper for questions.
1 paper · 0 benchmarks
Annotating data is a time-consuming and costly task, but it is inherently required for supervised machine learning.
1 paper · 0 benchmarks
To explore the nascent area of sustainable venture capital, a review of related research was conducted and social entrepreneurs & investors interviewed to construct a questionnaire assessing the interests and intentions of current & future…
1 paper · 0 benchmarks
Table-ACM12K (TACM12K) is a relational table dataset derived from the ACM heterogeneous graph dataset.
1 paper · 1 benchmark
TERRA-REF (TERRA-REF, An open reference data set from high resolution genomics, phenomics, and imaging sensors)
The ARPA-E funded TERRA-REF project is generating open-access reference datasets for the study of plant sensing, genomics, and phenomics.
1 paper · 0 benchmarks
Table-LastFm2K (TLF2K) is a relational table dataset derived from the classical LastFM2K dataset.
1 paper · 1 benchmark
TML1M (Table-MovieLens1M)
Table-MovieLens1M (TML1M) is a relational table dataset derived from the classical MovieLens1M dataset.
1 paper · 1 benchmark
TeleSim (TeleSim: A Network-Aware Testbed and Benchmark Dataset for Telerobotic Applications)
TeleSim is a network-aware hardware-in-the-loop dataset designed to evaluate the performance of telerobotic systems under varying network conditions.
1 paper · 0 benchmarks
The Reddit Climate Change Dataset is a dataset of 620K Reddit posts and 4.6M comments - all mentions of the terms "climate" and "change" until 2022-09-01 across the entire Reddit social network.
1 paper · 0 benchmarks
TimeGraph (TimeGraph: Synthetic Benchmark Datasets for Robust Time-Series Causal Discovery)
TimeGraph is a comprehensive suite of synthetic datasets designed to benchmark causal discovery algorithms on time-series data.
1 paper · 0 benchmarks
This repository contains data for the NeurIPS conference paper titled "Harnessing Machine Learning for Single-Shot Measurement of Free Electron Laser Pulse Power".
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
We introduce a dataset consisting of 1314 samples, including users’ tweets and bios.
1 paper · 0 benchmarks
AI-based digital twins are at the leading edge of theIndustry 4.0 revolution, which are technologically empowered bythe Internet of Things and real-time data analysis.
1 paper · 0 benchmarks
This data contains the election polls for the 2004, 2008, 2012, and 2016 US presidential election by state including data on undecided voter proportions.
1 paper · 0 benchmarks
Uniswap (Replication Data for: Uniswap Daily Transaction Indices by Network)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The prospective upper body thermal images SARS-CoV2 association study was designed to test the hypothesis that thermal videos may aid in the early diagnosis of COVID-19.
1 paper · 0 benchmarks
Introduction This dataset was gathered during the Vid2RealHRI study of humans’ perception of robots' intelligence in the context of an incidental Human-Robot encounter.
1 paper · 0 benchmarks
Context of the data sets The Zooniverse platform (www.zooniverse.org) has successfully built a large community of volunteers contributing to citizen science projects.
1 paper · 0 benchmarks
Enriched Voxceleb speakers' data of 1,715 celebrities with height gathered from Wikidata Other columns are taken from voxcelebenrichmentagegender
1 paper · 0 benchmarks
WDC Block is a benchmark for comparing the performance of blocking methods that are used as part of entity resolution pipelines.
1 paper · 3 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.