{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/quantifying-interpretability-and-trust-in","title":"Quantifying Interpretability and Trust in Machine Learning Systems","arxiv_id":"1901.08558","date":"2019-01-20","proceeding":null,"authors":["Philipp Schmidt","Felix Biessmann"],"abstract":"Decisions by Machine Learning (ML) models have become ubiquitous. Trusting\nthese decisions requires understanding how algorithms take them. Hence\ninterpretability methods for ML are an active focus of research. A central\nproblem in this context is that both the quality of interpretability methods as\nwell as trust in ML predictions are difficult to measure. Yet evaluations,\ncomparisons and improvements of trust and interpretability require quantifiable\nmeasures. Here we propose a quantitative measure for the quality of\ninterpretability methods. Based on that we derive a quantitative measure of\ntrust in ML decisions. Building on previous work we propose to measure\nintuitive understanding of algorithmic decisions using the information transfer\nrate at which humans replicate ML model predictions. We provide empirical\nevidence from crowdsourcing experiments that the proposed metric robustly\ndifferentiates interpretability methods. The proposed metric also demonstrates\nthe value of interpretability for ML assisted human decision making: in our\nexperiments providing explanations more than doubled productivity in annotation\ntasks. However unbiased human judgement is critical for doctors, judges, policy\nmakers and others. Here we derive a trust metric that identifies when human\ndecisions are overly biased towards ML predictions. Our results complement\nexisting qualitative work on trust and interpretability by quantifiable\nmeasures that can serve as objectives for further improving methods in this\nfield of research.","url_abs":"http://arxiv.org/abs/1901.08558v1","url_pdf":"http://arxiv.org/pdf/1901.08558v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"quantifying-interpretability-and-trust-in","repo_url":"https://github.com/anahid1988/DeepRUL","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"},{"task_slug":"decision-making","task_name":"Decision Making"}],"methods":[{"method_slug":"interpretability","method_name":"Interpretability"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1901.08558","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}