{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/tabarena-a-living-benchmark-for-machine","title":"TabArena: A Living Benchmark for Machine Learning on Tabular Data","arxiv_id":"2506.16791","date":"2025-06-20","proceeding":null,"authors":["Nick Erickson","Lennart Purucker","Andrej Tschalzev","David Holzmüller","Prateek Mutalik Desai","David Salinas","Frank Hutter"],"abstract":"With the growing popularity of deep learning and foundation models for tabular data, the need for standardized and reliable benchmarks is higher than ever. However, current benchmarks are static. Their design is not updated even if flaws are discovered, model versions are updated, or new models are released. To address this, we introduce TabArena, the first continuously maintained living tabular benchmarking system. To launch TabArena, we manually curate a representative collection of datasets and well-implemented models, conduct a large-scale benchmarking study to initialize a public leaderboard, and assemble a team of experienced maintainers. Our results highlight the influence of validation method and ensembling of hyperparameter configurations to benchmark models at their full potential. While gradient-boosted trees are still strong contenders on practical tabular datasets, we observe that deep learning methods have caught up under larger time budgets with ensembling. At the same time, foundation models excel on smaller datasets. Finally, we show that ensembles across models advance the state-of-the-art in tabular machine learning and investigate the contributions of individual models. We launch TabArena with a public leaderboard, reproducible code, and maintenance protocols to create a living benchmark available at https://tabarena.ai.","url_abs":"https://arxiv.org/abs/2506.16791v2","url_pdf":"https://arxiv.org/pdf/2506.16791v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"tabarena-a-living-benchmark-for-machine","repo_url":"https://github.com/autogluon/tabrepo","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2506.16791","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2506.16791"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/autogluon/tabrepo","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"unverified":3},"by_repo_kind":{"listed":{"samples":3,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"60dec2a6f22b44f4","entry":"compute_winrate","repo":"autogluon/tabrepo","repo_kind":"listed","path":"packages/bencheval/src/bencheval/winrate_utils.py","file_url":"https://github.com/autogluon/tabrepo/blob/HEAD/packages/bencheval/src/bencheval/winrate_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"60dec2a6f22b44f4"}},{"code_sha256_prefix":"610daec8cce6e85f","entry":"compute_winrate_matrix","repo":"autogluon/tabrepo","repo_kind":"listed","path":"packages/bencheval/src/bencheval/winrate_utils.py","file_url":"https://github.com/autogluon/tabrepo/blob/HEAD/packages/bencheval/src/bencheval/winrate_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"610daec8cce6e85f"}},{"code_sha256_prefix":"f8322a48a2fa3ae9","entry":"get_bootstrap_result_lst","repo":"autogluon/tabrepo","repo_kind":"listed","path":"packages/bencheval/src/bencheval/evaluator.py","file_url":"https://github.com/autogluon/tabrepo/blob/HEAD/packages/bencheval/src/bencheval/evaluator.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"f8322a48a2fa3ae9"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}