Home › Datasets › modality › Tabular

Tabular datasets

archive 2025-07-28

267 datasets carry the modality tag "Tabular", ordered by the archive's paper count. Page 3 of 6: 48 shown of 267. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Tabular datasets 97–144 of 267

The dataset represents data generated from a commonly used model in population genetics.
2 papers · 0 benchmarks
TAP (Traffic Accident Prediction data repository)
The Traffic Accident Prediction (TAP) data repository offers extensive coverage for 1,000 US cities (TAP-city) and 49 states (TAP-state), providing real-world road structure data that can be easily used for graph-based machine learning…
2 papers · 0 benchmarks
TNCR Dataset (Table Net Detection and Classification Dataset)
We present TNCR, a new table dataset with varying image quality collected from free open source websites.
2 papers · 0 benchmarks
Titanic (Titanic - Machine Learning from Disaster)
Titanic Dataset Description Overview The data is divided into two groups: - Training set (train.csv): Used to build machine learning models.
2 papers · 1 benchmark
UPenn-GBM (The University of Pennsylvania glioblastoma (UPenn-GBM) cohort)
This collection comprises multi-parametric magnetic resonance imaging (mpMRI) scans for de novo Glioblastoma (GBM) patients from the University of Pennsylvania Health System, coupled with patient demographics, clinical outcome (e.g.,…
2 papers · 0 benchmarks
The code to create the dataset is available here.
2 papers · 2 benchmarks
WDC SOTAB is a benchmark that features two annotation tasks: Column Type Annotation and Columns Property Annotation.
2 papers · 2 benchmarks
WikiTableSet (Wikipedia Table Image Dataset)
WikiTableSet is a large publicly available image-based table recognition dataset in three languages built from Wikipedia.
2 papers · 0 benchmarks
Wyze Rule Recommendation Dataset.
2 papers · 0 benchmarks
bcTCGA (The Cancer Genome Atlas Program)
This data set comes from breast cancer tissue samples deposited to The Cancer Genome Atlas (TCGA) project.
2 papers · 0 benchmarks
e2006 (10-K Corpus)
From the official description: > The corpus contains 10-K reports from many US companies during years > 1996-2006, as well as measured volatility of stock returns for the > twelve-month periods preceding and following each report.
2 papers · 0 benchmarks
kickstarter (Funding Successful Projects on Kickstarter)
Kickstarter is a community of more than 10 million people comprising of creative, tech enthusiasts who help in bringing creative project to life.
2 papers · 1 benchmark
news20 (NewsWeeder: learning to filter netnews)
Two datasets featuring binary and multi-class classification.
2 papers · 0 benchmarks
The datasets used and analysed from the glucose clamp study are available in this DIF file.
1 paper · 0 benchmarks
The datasets used and analysed from the glucose clamp study are available in this Excel file.
1 paper · 0 benchmarks
5DOF GB Interpolation (Five Degree-of-Freedom Grain Boundary Interpolation)
These are larger MATLAB .mat files required for reproducing plots from the sgbaird-5DOF/interp repository for grain boundary property interpolation.
1 paper · 0 benchmarks
This paper constructs 7-digit product Supply-Use Tables (SUTs) and symmetric Input-Output Tables (IOTs) for the Indian economy using microdata from the Annual Survey of Industries (ASI) for the period 2016-2021.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Multimodal object recognition is still an emerging field.
1 paper · 0 benchmarks
AgroLens (AgroLens Soil Prediction Dataset)
This dataset has been curated for a student research project at the Technische Hochschule Ingolstadt with Mi4Poeople and its Soil project (https://de.mi4people.org/soil-quality-evaluation-system).
1 paper · 0 benchmarks
This R package, documented in a very similar way to the book R4DS, provides functions to replicate the original Stata results from the book An Advanced Guide to Trade Policy Analysis.
1 paper · 0 benchmarks
Data collected from two budget surveys (FY2021 in 2020 and FY2022 in 2021) in collaboration with the City of Austin budget department.
1 paper · 0 benchmarks
BASIR (BASIR_Budget_Assisted_Sectoral_Impact_Ranking)
Government fiscal policies, particularly annual union budgets, exert significant influence on financial markets.
1 paper · 0 benchmarks
BODMAS (Blue Hexagon Open Dataset for Malware AnalysiS)
We collaborate with Blue Hexagon to release a dataset containing timestamped malware samples and well-curated family information for research purposes.
1 paper · 0 benchmarks
Hand-disambiguation of a sample of U.S.
1 paper · 0 benchmarks
BreastRates4 ([MIMBCD-UI] UTA4: Rates Dataset)
Several datasets are fostering innovation in higher-level functions for everyone, everywhere.
1 paper · 0 benchmarks
CANDOR Corpus (CANDOR = Conversation: A Naturalistic Dataset of Online Recordings)
The CANDOR corpus is a large, novel, multimodal corpus of 1,656 recorded conversations in spoken English.
1 paper · 0 benchmarks
CSTS (Correlation Structures in Time Series)
CSTS: Correlation Structures in Time Series CSTS is a comprehensive synthetic benchmarking dataset designed specifically for evaluating correlation structure discovery in time series data.
1 paper · 0 benchmarks
CVR (Congressional Voting Records Data Set)
This data set includes votes for each of the U.S.
1 paper · 1 benchmark
This dataset contains a two-column CSV file, where the first column ("ValidcitingDOI") contains the DOI of a citing entity retrieved in Crossref, while the second column ("InvalidcitedDOI") contains the invalid DOI of a cited entity…
1 paper · 0 benchmarks
Co/FeMn bilayers measured.
1 paper · 0 benchmarks
Coastal Inundation Maps with Floodwater Depth Values (Simulated Flood Inundation Maps of Abu Dhabi's Coast Under Different Shoreline Protection Scenarios)
This dataset provides simulated flood inundation maps of Abu Dhabi's coast under 174 different shoreline protection scenarios.
1 paper · 1 benchmark
CompMix-IR Dataset Overview: Characteristics: CompMix-IR is a heterogeneous knowledge retrieval benchmark dataset, featuring four knowledge types (text, knowledge graphs, tables, and infoboxes), 9,400+ QA pairs, and a corpus of 10 million…
1 paper · 0 benchmarks
In everyday language processing, sentence context affects how readers and listeners process upcoming words.
1 paper · 0 benchmarks
Concrete is the most important material in civil engineering.
1 paper · 1 benchmark
The dataset contains 30 million cryptocurrency-related tweets from 10.10.2020 to 3.3.2021.
1 paper · 0 benchmarks
DBFC Dataset (Single Direct Borohydride Fuel Cell Dataset)
This dataset includes Direct Borohydride Fuel Cell (DBFC) impedance and polarization test in anode with Pd/C, Pt/C and Pd decorated Ni–Co/rGO catalysts.
1 paper · 0 benchmarks
[comment]:<> (Data for the paper "Deciphering Environmental Air Pollution with Large Scale City Data") Main Dataset citypollutiondata.csv Relevant Columns: Date: Date of the sample City: City of the sample Xmedian: Median value of the…
1 paper · 0 benchmarks
IOPS and Latency measurements of a real data storage system
1 paper · 0 benchmarks
Overview This dataset was collected during a pilot study that evaluated the virtual Cross Array Task (CAT) platform as an assessment tool for algorithmic thinking (AT) skills among K-12 students in Swiss compulsory education.
1 paper · 0 benchmarks
Overview of the scoping review paper corpus, sorted by their diferent intent types, categories, and subcategories.
1 paper · 0 benchmarks
The dataset is generated from the study of computational reproducibility of Jupyter notebooks from biomedical publications.
1 paper · 0 benchmarks
This repository contains the dataset for the study of the computational reproducibility of Jupyter notebooks from biomedical publications.
1 paper · 0 benchmarks
Collected data from two distinct experiments in immersive, interactive VR where participants performed dynamic tasks as their eye, head, and hand movements were recorded.
1 paper · 0 benchmarks
DocBank-TB (DocBank-Table)
This dataset consisting 500 set of caption, table and coresponding paper page, processed from DocBank.
1 paper · 0 benchmarks
The data used for all results in this paper can be found here.
1 paper · 0 benchmarks
ELMTEX Dataset (ELMTEX Dataset: Fine-Tuning Large Language Models for Structured Clinical Information Extraction)
We introduced a new dataset of clinical report summaries, annotated with structured information across 15 categories.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.