{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/amlb-an-automl-benchmark","title":"AMLB: an AutoML Benchmark","arxiv_id":"2207.12560","date":"2022-07-25","proceeding":null,"authors":["Pieter Gijsbers","Marcos L. P. Bueno","Stefan Coors","Erin LeDell","Sébastien Poirier","Janek Thomas","Bernd Bischl","Joaquin Vanschoren"],"abstract":"Comparing different AutoML frameworks is notoriously challenging and often done incorrectly. We introduce an open and extensible benchmark that follows best practices and avoids common mistakes when comparing AutoML frameworks. We conduct a thorough comparison of 9 well-known AutoML frameworks across 71 classification and 33 regression tasks. The differences between the AutoML frameworks are explored with a multi-faceted analysis, evaluating model accuracy, its trade-offs with inference time, and framework failures. We also use Bradley-Terry trees to discover subsets of tasks where the relative AutoML framework rankings differ. The benchmark comes with an open-source tool that integrates with many AutoML frameworks and automates the empirical evaluation process end-to-end: from framework installation and resource allocation to in-depth evaluation. The benchmark uses public data sets, can be easily extended with other AutoML frameworks and tasks, and has a website with up-to-date results.","url_abs":"https://arxiv.org/abs/2207.12560v2","url_pdf":"https://arxiv.org/pdf/2207.12560v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"amlb-an-automl-benchmark","repo_url":"https://github.com/openml/automlbenchmark","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"amlb-an-automl-benchmark","repo_url":"https://github.com/shchur/automlbenchmark","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"automl","task_name":"AutoML"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2207.12560","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2207.12560"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/shchur/automlbenchmark","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/openml/automlbenchmark","reach":null}],"summary":{"ran_draft_wrong":3,"unverified":3},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1},"listed":{"samples":3,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"420471e1e7b46a85","entry":"generate_framework_gallery","repo":"openml/automlbenchmark","repo_kind":"official","path":"docs/website/scripts/generate_index.py","file_url":"https://github.com/openml/automlbenchmark/blob/HEAD/docs/website/scripts/generate_index.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"420471e1e7b46a85"}},{"code_sha256_prefix":"17322beb8b543527","entry":"load_framework_definitions","repo":"openml/automlbenchmark","repo_kind":"official","path":"docs/website/scripts/generate_index.py","file_url":"https://github.com/openml/automlbenchmark/blob/HEAD/docs/website/scripts/generate_index.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"17322beb8b543527"}},{"code_sha256_prefix":"8f8bcbb6f947b63b","entry":"parse_frameworks","repo":"openml/automlbenchmark","repo_kind":"official","path":"docs/website/scripts/generate_index.py","file_url":"https://github.com/openml/automlbenchmark/blob/HEAD/docs/website/scripts/generate_index.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"8f8bcbb6f947b63b"}},{"code_sha256_prefix":"3d2c07f3353df4d6","entry":"is_data_frame","repo":"shchur/automlbenchmark","repo_kind":"listed","path":"amlb/datautils.py","file_url":"https://github.com/shchur/automlbenchmark/blob/HEAD/amlb/datautils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3d2c07f3353df4d6"}},{"code_sha256_prefix":"fedce06bb6f8cf4a","entry":"read_csv","repo":"shchur/automlbenchmark","repo_kind":"listed","path":"amlb/datautils.py","file_url":"https://github.com/shchur/automlbenchmark/blob/HEAD/amlb/datautils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"fedce06bb6f8cf4a"}},{"code_sha256_prefix":"f9fb6e042d6be804","entry":"setup","repo":"shchur/automlbenchmark","repo_kind":"listed","path":"amlb/logger.py","file_url":"https://github.com/shchur/automlbenchmark/blob/HEAD/amlb/logger.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f9fb6e042d6be804"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}