Home › Datasets › modality › Tables
Tables datasets
archive 2025-07-28
50 datasets carry the modality tag "Tables", ordered by the archive's paper count. Page 1 of 2: 48 shown of 50. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Tables datasets 1–48 of 50
TAT-QA (Tabular And Textual dataset for Question Answering) is a large-scale QA dataset, aiming to stimulate progress of QA research over more complex and realistic tabular and textual data, especially those requiring numerical reasoning.
76 papers · 1 benchmark
WebSRC (WebSRC: A Dataset for Web-Based Structural Reading Comprehension)
WebSRC is a novel Web-based Structural Reading Comprehension dataset.
22 papers · 2 benchmarks
GitTables is a corpus of currently 1M relational tables extracted from CSV files in GitHub covering 96 topics.
16 papers · 0 benchmarks
SinD (A Drone Dataset at Signalized Intersection in China)
The SIND dataset is based on 4K video captured by drones, providing information including traffic participant trajectories, traffic light status, and high-definition maps
14 papers · 0 benchmarks
VNAT (VPN/NONVPN NETWORK APPLICATION TRAFFIC DATASET)
This dataset is a collection of labelled PCAP files, both encrypted and unencrypted, across 10 applications, as well as a pandas dataframe in HDF5 format containing detailed metadata summarizing the connections from those files.
5 papers · 0 benchmarks
The M5Product dataset is a large-scale multi-modal pre-training dataset with coarse and fine-grained annotations for E-products.
4 papers · 0 benchmarks
eICU-CRD (eICU Collaborative Research Database)
The eICU Collaborative Research Database is a large multi-center critical care database made available by Philips Healthcare in partnership with the MIT Laboratory for Computational Physiology.
4 papers · 2 benchmarks
MMCode is a multi-modal code generation dataset designed to evaluate the problem-solving skills of code language models in visually rich contexts (i.e.
3 papers · 0 benchmarks
SKAB (Skoltech Anomaly Benchmark)
SKAB is designed for evaluating algorithms for anomaly detection.
3 papers · 2 benchmarks
The ArxivPapers dataset is an unlabelled collection of over 104K papers related to machine learning and published on arXiv.org between 2007–2020.
2 papers · 0 benchmarks
This repository contains the database of the FEM simulation of axially impacted various configurations of the square crash boxes.
2 papers · 0 benchmarks
GIRT-Data (GitHub Issue Report Template Dataset)
GIRT-Data is the first and largest dataset of issue report templates (IRTs) in both YAML and Markdown format.
2 papers · 0 benchmarks
This dataset presents a set of large-scale ridesharing Dial-a-Ride Problem (DARP) instances.
2 papers · 0 benchmarks
The original dataset was provided by Orange telecom in France, which contains anonymized and aggregated human mobility data.
2 papers · 0 benchmarks
Articles originating from subreddits with explicitly stated ideologies are categorized into three groups: 72,488 articles in the Liberal class, 79,573 articles in the Conservative class, and 225,083 articles in the Restricted class.
2 papers · 1 benchmark
The SegmentedTables dataset is a collection of almost 2,000 tables extracted from 352 machine learning papers.
2 papers · 0 benchmarks
The SheetCopilot dataset contains 28 evaluation workbooks and 221 spreadsheet manipulation tasks that are applied to these workbooks.
2 papers · 1 benchmark
ASRD (Anime Style Recognition Dataset)
A well-labeled challenging dataset, to facilitate the research on style recognition on anime images by collecting images from 190 anime and cartoon works covering 93 years from 13 countries and regions, 2D and 3D work into consideration…
1 paper · 0 benchmarks
Yavuz Selim TASPINAR, Murat KOKLU and Mustafa ALTIN Citation Request : 1: KOKLU M., TASPINAR Y.S., (2021).
1 paper · 0 benchmarks
This dataset includes Direct Borohydride Fuel Cell (DBFC) impedance and polarization test in anode with Pd/C, Pt/C and Pd decorated Ni–Co/rGO catalysts.
1 paper · 0 benchmarks
IOPS and Latency measurements of a real data storage system
1 paper · 0 benchmarks
The dataset is generated from the study of computational reproducibility of Jupyter notebooks from biomedical publications.
1 paper · 0 benchmarks
This repository contains the dataset for the study of the computational reproducibility of Jupyter notebooks from biomedical publications.
1 paper · 0 benchmarks
The dataset contains two Pareto-fronts: - The Pareto-front for the 2-objective problem - The Pareto-front for the 3-objective problem Each Pareto-front contains a set of points, with coordinates given by their objectives.
1 paper · 0 benchmarks
This is a real-world industrial benchmark dataset from a major medical device manufacturer for the prediction of customer escalations.
1 paper · 0 benchmarks
This dataset contains pre-processed versions of datasets introduced in prior works.
1 paper · 0 benchmarks
We present a comprehensive dataset comprising a vast collection of raw mineral samples for the purpose of mineral recognition.
1 paper · 0 benchmarks
US Macroeconomic dataset containing 14 time series of monthly observations.
1 paper · 0 benchmarks
This dataset artifact contains the intermediate datasets from pipeline executions necessary to reproduce the results of the paper.
1 paper · 0 benchmarks
This dataset are about Nafion 112 membrane standard tests and MEA activation tests of PEM fuel cell in various operation condition.
1 paper · 0 benchmarks
The Perfume Co-Preference Network dataset comprises comprehensive user reviews and ratings collected from the Persian retail platform Atrafshan.
1 paper · 0 benchmarks
This repository contains a dataset and machine learning algorithms to detect poisoned water from clean water via using equivalent Smartphone embedded Wi-Fi CSI data.
1 paper · 0 benchmarks
We conducted a large crowdsourcing study of click patterns in an interactive segmentation scenario and collected 475K real-user clicks.
1 paper · 0 benchmarks
The dataset provides information about 450 HYIPs collected between November 2020 and September 2021.
1 paper · 0 benchmarks
ata Set Name: Rice Dataset (Commeo and Osmancik) Abstract: A total of 3810 rice grain's images were taken for the two species (Cammeo and Osmancik), processed and feature inferences were made.
1 paper · 0 benchmarks
SCG (SCG Dataset from Graph Neural Networks in Supply Chain Analytics and Optimization: Concepts, Perspectives, Dataset & Benchmarks)
Abstract: Graph Neural Networks (GNNs) have recently gained traction in transportation, bioinformatics, language and image processing, but research on their application to supply chain management remains limited.
1 paper · 1 benchmark
Our SRSD (Feynman) datasets are designed to discuss the performance of Symbolic Regression for Scientific Discovery.
1 paper · 0 benchmarks
Our SRSD (Feynman) datasets are designed to discuss the performance of Symbolic Regression for Scientific Discovery.
1 paper · 0 benchmarks
Our SRSD (Feynman) datasets are designed to discuss the performance of Symbolic Regression for Scientific Discovery.
1 paper · 0 benchmarks
Dataset Card for SemTabNet This dataset accompanies the following paper: Title: Statements: Universal Information Extraction from Tables with Large Language Models for ESG KPIs Authors: Lokesh Mishra, Sohayl Dhibi, Yusik Kim, Cesar…
1 paper · 1 benchmark
> The StatCan Dialogue Dataset: Retrieving Data Tables through Conversations with Genuine Intents > > Xing Han Lu, Siva Reddy, Harm de Vries > > EACL 2023 | | | | | | | :--: | :--: | :--: | :--: | :--: | | Code | Huggingface | Request on…
1 paper · 1 benchmark
We introduce a dataset consisting of 1314 samples, including users’ tweets and bios.
1 paper · 0 benchmarks
This dataset collection includes three files used for the experiments.
1 paper · 0 benchmarks
It contains data from two different realities: Food.com, a well-known American recipe site, and Planeat, an Italian site that allows you to plan recipes to save food waste.
1 paper · 0 benchmarks
bSDD (buildingSMART Data Dictionary)
The buildingSMART Data Dictionary (bSDD) is an online service that hosts classifications and their properties, allowed values, units and translations.
1 paper · 0 benchmarks
data_qe (Federal Reserve Quantitative Easing Data)
This file contains the data and code for the publication "The Federal Reserve's Response to the Global Financial Crisis and Its Long-Term Impact: An Interrupted Time-Series Natural Experimental Analysis" by A.
1 paper · 0 benchmarks
It is a competition on kaggle with stroke Prediction, which is heavily imbalanced.
1 paper · 0 benchmarks
Item-wise accuracies in six benchmarks from Open LLM Leaderboard 1 scraped from huggingface.co and used for metabench analyses and construction.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.