Home › Datasets › modality › Tabular
Tabular datasets
archive 2025-07-28
267 datasets carry the modality tag "Tabular", ordered by the archive's paper count. Page 4 of 6: 48 shown of 267. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Tabular datasets 145–192 of 267
EUCA dataset description Associated Paper: EUCA: the End-User-Centered Explainable AI Framework Authors: Weina Jin, Jianyu Fan, Diane Gromala, Philippe Pasquier, Ghassan Hamarneh Introduction: EUCA dataset is for modelling personalized or…
1 paper · 0 benchmarks
EUEN17037 Daylight and View Standard Test Dataset.
1 paper · 0 benchmarks
The EVI dataset is a challenging, multilingual spoken-dialogue dataset with 5,506 dialogues in English, Polish, and French.
1 paper · 3 benchmarks
Each HDF5 file has the following structure: energy Dataset {100000, 1} layer0 Dataset {100000, 3, 96} layer1 Dataset {100000, 12, 12} layer2 Dataset {100000, 12, 6} overflow Dataset {100000, 3} In practice, each file is a collection of…
1 paper · 0 benchmarks
Provide: a high-level explanation of the dataset characteristics explain motivations and summary of its content potential use cases of the dataset Collection of Error Grid data files.
1 paper · 0 benchmarks
The EyeInfo Dataset is an open-source eye-tracking dataset created by Fabricio Batista Narcizo, a research scientist at the IT University of Copenhagen (ITU) and GN Audio A/S (Jabra), Denmark.
1 paper · 0 benchmarks
Tables of the blendshapes from a group of the images of the FER2013 dataset, generated using MediaPipe library, based on the ARKit face blendshapes.
1 paper · 0 benchmarks
Optical images of printed circuit boards as well as detailed annotations of any text, logos, and surface-mount devices (SMDs).
1 paper · 0 benchmarks
DOI: https://doi.org/10.7910/DVN/O4CRXK The most comprehensive standardised data on Malaysian federal and state elections from 1955 to the present.
1 paper · 0 benchmarks
This dataset contains the publication data underlying the French Open Science Monitor.
1 paper · 0 benchmarks
GMSC (Give Me Some Credit)
Data for a Kaggle competition Banks play a crucial role in market economies.
1 paper · 0 benchmarks
This is the static test data from the study "Global Geolocated Realtime Data of Interfleet Urban Transit Bus Iding" collected by GRD-TRT-BUF-4I.
1 paper · 0 benchmarks
GenoTEX (An LLM Agent Benchmark for Automated Gene Expression Data Analysis)
GenoTEX (Genomics Data Automatic Exploration Benchmark) is a benchmark dataset for the automated analysis of gene expression data to identify disease-associated genes while considering the influence of other biological factors.
1 paper · 0 benchmarks
Genre2Movies (Compositional queries for Movie recommendation)
Genre annotations for movies The file genre2movies.csv contains genre-movie tuples based on Wikidata annotations (https://www.wikidata.org/).
1 paper · 0 benchmarks
Dataset Description This dataset contains rental property listings scraped from Tonaton.com, one of Ghana's leading online classifieds platforms.
1 paper · 0 benchmarks
The National Health and Nutrition Examination Survey (NHANES) provides data on the health and environmental exposure of the non-institutionalized US population.
1 paper · 0 benchmarks
Inpatient claims, Outpatient claims and Beneficiary details of each provider.
1 paper · 1 benchmark
Heteroatom doped graphene supercapacitor feature data is gathered from various literatures for use in machine learning tasks.
1 paper · 0 benchmarks
Here the dataset described in Hitchhiking Rides Dataset: Two decades of crowd-sourced records on stochastic traveling(https://arxiv.org/abs/2506.21946) is published.
1 paper · 0 benchmarks
A maintained database tracks ICLR submissions and reviews, augmented with author profiles and higher-level textual features.
1 paper · 0 benchmarks
IEIs (Ion and Electron Insulators)
We would like to introduce three types of ion and electron insulators, i.e.
1 paper · 0 benchmarks
We present two multi-modal datasets, one for Main Board IPOs, and the other for Small and Medium Enterprises (SME) IPOs.
1 paper · 0 benchmarks
Context The Kepler Space Observatory is a NASA-build satellite that was launched in 2009.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Dataset is a CSV file, that contains evaluation scores given by a panel of LLMs to responses produced by other LLMs .
1 paper · 0 benchmarks
LimeSoDa (Precision Liming Soil Datasets)
Precision Liming Soil Datasets (LimeSoDa) is a collection of 31 datasets from a field- and farm-scale soil mapping context.
1 paper · 0 benchmarks
The LinkedResults dataset contains around 1,600 results capturing performance of machine learning models from tables of 239 papers.
1 paper · 0 benchmarks
CSV file with a list of all examined OWL reasoners.
1 paper · 0 benchmarks
The LoRA Weight Size Evaluation (LoRA-WiSE) is a comprehensive benchmark specifically designed to evaluate LoRA dataset size recovery methods for generative models LoRA-WiSE spans various dataset sizes, backbones, ranks, and…
1 paper · 0 benchmarks
This dataset contains pre-processed versions of datasets introduced in prior works.
1 paper · 0 benchmarks
MIMI dataset (Multi-aspect Integrated Migration Indicators dataset)
Nowadays, new branches of research are proposing the use of non-traditional data sources for the study of migration trends in order to find an original methodology to answer open questions about cross-border human mobility.
1 paper · 0 benchmarks
MPOSE2021 (MPOSE2021 Dataset for Short-time Human Action Recognition)
MPOSE2021, a dataset for real-time short-time HAR, suitable for both pose-based and RGB-based methodologies.
1 paper · 0 benchmarks
MVX incorporates realistic physical world simulation with a differentiable accurate ray tracing wireless simulation that includes multi-agent and multimodal datasets for AI-driven digital twin applications in vehicular communication…
1 paper · 1 benchmark
This dataset was developed within an analysis of research data generated and managed within the University of Bologna, with respect to the differences and commonalities between disciplines and potential challenges for institutional data…
1 paper · 0 benchmarks
A new in-context visual question answering dataset encompassing interleaved image and EHR data derived from MIMIC-IV and MIMIC-CXR-JPG databases.
1 paper · 0 benchmarks
This dataset contains demographic and personal health information for individuals, along with the corresponding medical insurance charges billed to them.
1 paper · 1 benchmark
MiMIC (Multi-Modal Indian Earnings Calls Dataset)
Predicting stock market prices following corporate earnings calls remains a significant challenge for investors and researchers alike, requiring innovative approaches that can process diverse information sources.
1 paper · 0 benchmarks
This dataset is a multi-labelled SMILES odor dataset with 138 odor descriptors.
1 paper · 1 benchmark
A detailed description of this dataset can be found in the Zenodo repository: https://zenodo.org/record/7845311#.ZK-jty9BxhE
1 paper · 0 benchmarks
A detailed description of this dataset can be found in the Zenodo repository: https://zenodo.org/record/7845361#.ZK-k7y9BxhE
1 paper · 0 benchmarks
A detailed description of this dataset can be found in the Zenodo repository: https://zenodo.org/record/8119042#.ZK-jJC9BxhE
1 paper · 0 benchmarks
NBA_Box_Scores_Odds (NBA Team-Level Box Score Statistics (2015-2019), Historical Win Percentages (2014-2018) and Betting Odds (2018/2019))
Dataset Description: NBA Team Statistics, Historical Performance & Betting Odds (2015-2019) Overview This dataset contains team-level box score statistics, historical win percentages, and closing betting odds for NBA games from 2015 to…
1 paper · 0 benchmarks
The NVALT-11 study considered the effect of profylactic brain radiation versus observation in (m=174) patients with advanced non-small cell lung cancer.
1 paper · 0 benchmarks
Te NVALT-8 study (m=200 participants) examined if nadroparin combined with chemotherapy could reduce cancer relapse after surgical removal of a non-small cell lung tumour.
1 paper · 0 benchmarks
ODDS (Outlier Detection DataSets (ODDS))
Outliers or anomalies are instances that do not conform to the norm of a dataset.
1 paper · 1 benchmark
OPFLearnData (OPFLearnData: Dataset for Learning AC Optimal Power Flow)
The datasets are resulting from OPFLearn.jl, a Julia package for creating AC OPF datasets.
1 paper · 0 benchmarks
The OTTO session dataset is a large-scale dataset intended for multi-objective recommendation research.
1 paper · 0 benchmarks
PDFM Embeddings are condensed vector representations designed to encapsulate the complex, multidimensional interactions among human behaviors, environmental factors, and local contexts at specific locations.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.