{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/automating-biomedical-data-science-through","title":"Automating biomedical data science through tree-based pipeline optimization","arxiv_id":"1601.07925","date":"2016-01-28","proceeding":null,"authors":["Randal S. Olson","Ryan J. Urbanowicz","Peter C. Andrews","Nicole A. Lavender","La Creis Kidd","Jason H. Moore"],"abstract":"Over the past decade, data science and machine learning has grown from a\nmysterious art form to a staple tool across a variety of fields in academia,\nbusiness, and government. In this paper, we introduce the concept of tree-based\npipeline optimization for automating one of the most tedious parts of machine\nlearning---pipeline design. We implement a Tree-based Pipeline Optimization\nTool (TPOT) and demonstrate its effectiveness on a series of simulated and\nreal-world genetic data sets. In particular, we show that TPOT can build\nmachine learning pipelines that achieve competitive classification accuracy and\ndiscover novel pipeline operators---such as synthetic feature\nconstructors---that significantly improve classification accuracy on these data\nsets. We also highlight the current challenges to pipeline optimization, such\nas the tendency to produce pipelines that overfit the data, and suggest future\nresearch paths to overcome these challenges. As such, this work represents an\nearly step toward fully automating machine learning pipeline design.","url_abs":"http://arxiv.org/abs/1601.07925v1","url_pdf":"http://arxiv.org/pdf/1601.07925v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"automating-biomedical-data-science-through","repo_url":"https://github.com/rhiever/tpot","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"hyperparameter-optimization","task_name":"Hyperparameter Optimization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}