{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/quantifying-contribution-and-propagation-of","title":"Quantifying contribution and propagation of error from computational steps, algorithms and hyperparameter choices in image classification pipelines","arxiv_id":"1903.00405","date":"2019-02-21","proceeding":null,"authors":["Aritra Chowdhury","Malik Magdon-Ismail","Bulent Yener"],"abstract":"Data science relies on pipelines that are organized in the form of\ninterdependent computational steps. Each step consists of various candidate\nalgorithms that maybe used for performing a particular function. Each algorithm\nconsists of several hyperparameters. Algorithms and hyperparameters must be\noptimized as a whole to produce the best performance. Typical machine learning\npipelines consist of complex algorithms in each of the steps. Not only is the\nselection process combinatorial, but it is also important to interpret and\nunderstand the pipelines. We propose a method to quantify the importance of\ndifferent components in the pipeline, by computing an error contribution\nrelative to an agnostic choice of computational steps, algorithms and\nhyperparameters. We also propose a methodology to quantify the propagation of\nerror from individual components of the pipeline with the help of a naive set\nof benchmark algorithms not involved in the pipeline. We demonstrate our\nmethodology on image classification pipelines. The agnostic and naive\nmethodologies quantify the error contribution and propagation respectively from\nthe computational steps, algorithms and hyperparameters in the image\nclassification pipeline. We show that algorithm selection and hyperparameter\noptimization methods like grid search, random search and Bayesian optimization\ncan be used to quantify the error contribution and propagation, and that random\nsearch is able to quantify them more accurately than Bayesian optimization.\nThis methodology can be used by domain experts to understand machine learning\nand data analysis pipelines in terms of their individual components, which can\nhelp in prioritizing different components of the pipeline.","url_abs":"http://arxiv.org/abs/1903.00405v1","url_pdf":"http://arxiv.org/pdf/1903.00405v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"quantifying-contribution-and-propagation-of","repo_url":"https://github.com/AriChow/error_propagation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"},{"task_slug":"bayesian-optimization","task_name":"Bayesian Optimization"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"hyperparameter-optimization","task_name":"Hyperparameter Optimization"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[{"method_slug":"random-search","method_name":"Random Search"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}