Home › Datasets › modality › Tabular
Tabular datasets
archive 2025-07-28
267 datasets carry the modality tag "Tabular", ordered by the archive's paper count. Page 3 of 6: 48 shown of 267. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Tabular datasets 97–144 of 267
The dataset represents data generated from a commonly used model in population genetics.
2 papers · 0 benchmarks
TAP (Traffic Accident Prediction data repository)
The Traffic Accident Prediction (TAP) data repository offers extensive coverage for 1,000 US cities (TAP-city) and 49 states (TAP-state), providing real-world road structure data that can be easily used for graph-based machine learning…
2 papers · 0 benchmarks
TNCR Dataset (Table Net Detection and Classification Dataset)
We present TNCR, a new table dataset with varying image quality collected from free open source websites.
2 papers · 0 benchmarks
Titanic (Titanic - Machine Learning from Disaster)
Titanic Dataset Description Overview The data is divided into two groups: - Training set (train.csv): Used to build machine learning models.
2 papers · 1 benchmark
UPenn-GBM (The University of Pennsylvania glioblastoma (UPenn-GBM) cohort)
This collection comprises multi-parametric magnetic resonance imaging (mpMRI) scans for de novo Glioblastoma (GBM) patients from the University of Pennsylvania Health System, coupled with patient demographics, clinical outcome (e.g.,…
2 papers · 0 benchmarks
The code to create the dataset is available here.
2 papers · 2 benchmarks
WDC SOTAB is a benchmark that features two annotation tasks: Column Type Annotation and Columns Property Annotation.
2 papers · 2 benchmarks
WikiTableSet is a large publicly available image-based table recognition dataset in three languages built from Wikipedia.
2 papers · 0 benchmarks
Wyze Rule Recommendation Dataset.
2 papers · 0 benchmarks
bcTCGA (The Cancer Genome Atlas Program)
This data set comes from breast cancer tissue samples deposited to The Cancer Genome Atlas (TCGA) project.
2 papers · 0 benchmarks
From the official description: > The corpus contains 10-K reports from many US companies during years > 1996-2006, as well as measured volatility of stock returns for the > twelve-month periods preceding and following each report.
2 papers · 0 benchmarks
kickstarter (Funding Successful Projects on Kickstarter)
Kickstarter is a community of more than 10 million people comprising of creative, tech enthusiasts who help in bringing creative project to life.
2 papers · 1 benchmark
news20 (NewsWeeder: learning to filter netnews)
Two datasets featuring binary and multi-class classification.
2 papers · 0 benchmarks
The datasets used and analysed from the glucose clamp study are available in this DIF file.
1 paper · 0 benchmarks
The datasets used and analysed from the glucose clamp study are available in this Excel file.
1 paper · 0 benchmarks
These are larger MATLAB .mat files required for reproducing plots from the sgbaird-5DOF/interp repository for grain boundary property interpolation.
1 paper · 0 benchmarks
This paper constructs 7-digit product Supply-Use Tables (SUTs) and symmetric Input-Output Tables (IOTs) for the Indian economy using microdata from the Annual Survey of Industries (ASI) for the period 2016-2021.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Multimodal object recognition is still an emerging field.
1 paper · 0 benchmarks
AgroLens (AgroLens Soil Prediction Dataset)
This dataset has been curated for a student research project at the Technische Hochschule Ingolstadt with Mi4Poeople and its Soil project (https://de.mi4people.org/soil-quality-evaluation-system).
1 paper · 0 benchmarks
This R package, documented in a very similar way to the book R4DS, provides functions to replicate the original Stata results from the book An Advanced Guide to Trade Policy Analysis.
1 paper · 0 benchmarks
Data collected from two budget surveys (FY2021 in 2020 and FY2022 in 2021) in collaboration with the City of Austin budget department.
1 paper · 0 benchmarks
BASIR (BASIR_Budget_Assisted_Sectoral_Impact_Ranking)
Government fiscal policies, particularly annual union budgets, exert significant influence on financial markets.
1 paper · 0 benchmarks
BODMAS (Blue Hexagon Open Dataset for Malware AnalysiS)
We collaborate with Blue Hexagon to release a dataset containing timestamped malware samples and well-curated family information for research purposes.
1 paper · 0 benchmarks
The dataset contains a total of 253,070 records, with 18 features.
1 paper · 0 benchmarks
Hand-disambiguation of a sample of U.S.
1 paper · 0 benchmarks
Several datasets are fostering innovation in higher-level functions for everyone, everywhere.
1 paper · 0 benchmarks
CANDOR Corpus (CANDOR = Conversation: A Naturalistic Dataset of Online Recordings)
The CANDOR corpus is a large, novel, multimodal corpus of 1,656 recorded conversations in spoken English.
1 paper · 0 benchmarks
CSTS (Correlation Structures in Time Series)
CSTS: Correlation Structures in Time Series CSTS is a comprehensive synthetic benchmarking dataset designed specifically for evaluating correlation structure discovery in time series data.
1 paper · 0 benchmarks
CVR (Congressional Voting Records Data Set)
This data set includes votes for each of the U.S.
1 paper · 1 benchmark
This dataset contains a two-column CSV file, where the first column ("ValidcitingDOI") contains the DOI of a citing entity retrieved in Crossref, while the second column ("InvalidcitedDOI") contains the invalid DOI of a cited entity…
1 paper · 0 benchmarks
Co/FeMn bilayers measured.
1 paper · 0 benchmarks
This dataset provides simulated flood inundation maps of Abu Dhabi's coast under 174 different shoreline protection scenarios.
1 paper · 1 benchmark
CompMix-IR Dataset Overview: Characteristics: CompMix-IR is a heterogeneous knowledge retrieval benchmark dataset, featuring four knowledge types (text, knowledge graphs, tables, and infoboxes), 9,400+ QA pairs, and a corpus of 10 million…
1 paper · 0 benchmarks
In everyday language processing, sentence context affects how readers and listeners process upcoming words.
1 paper · 0 benchmarks
Concrete is the most important material in civil engineering.
1 paper · 1 benchmark
The dataset contains 30 million cryptocurrency-related tweets from 10.10.2020 to 3.3.2021.
1 paper · 0 benchmarks
This dataset includes Direct Borohydride Fuel Cell (DBFC) impedance and polarization test in anode with Pd/C, Pt/C and Pd decorated Ni–Co/rGO catalysts.
1 paper · 0 benchmarks
[comment]:<> (Data for the paper "Deciphering Environmental Air Pollution with Large Scale City Data") Main Dataset citypollutiondata.csv Relevant Columns: Date: Date of the sample City: City of the sample Xmedian: Median value of the…
1 paper · 0 benchmarks
IOPS and Latency measurements of a real data storage system
1 paper · 0 benchmarks
Overview This dataset was collected during a pilot study that evaluated the virtual Cross Array Task (CAT) platform as an assessment tool for algorithmic thinking (AT) skills among K-12 students in Swiss compulsory education.
1 paper · 0 benchmarks
Overview of the scoping review paper corpus, sorted by their diferent intent types, categories, and subcategories.
1 paper · 0 benchmarks
The dataset is generated from the study of computational reproducibility of Jupyter notebooks from biomedical publications.
1 paper · 0 benchmarks
This repository contains the dataset for the study of the computational reproducibility of Jupyter notebooks from biomedical publications.
1 paper · 0 benchmarks
Collected data from two distinct experiments in immersive, interactive VR where participants performed dynamic tasks as their eye, head, and hand movements were recorded.
1 paper · 0 benchmarks
This dataset consisting 500 set of caption, table and coresponding paper page, processed from DocBank.
1 paper · 0 benchmarks
The data used for all results in this paper can be found here.
1 paper · 0 benchmarks
ELMTEX Dataset (ELMTEX Dataset: Fine-Tuning Large Language Models for Structured Clinical Information Extraction)
We introduced a new dataset of clinical report summaries, annotated with structured information across 15 categories.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.