{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/input-prioritization-for-testing-neural","title":"Input Prioritization for Testing Neural Networks","arxiv_id":"1901.03768","date":"2019-01-11","proceeding":null,"authors":["Taejoon Byun","Vaibhav Sharma","Abhishek Vijayakumar","Sanjai Rayadurgam","Darren Cofer"],"abstract":"Deep neural networks (DNNs) are increasingly being adopted for sensing and\ncontrol functions in a variety of safety and mission-critical systems such as\nself-driving cars, autonomous air vehicles, medical diagnostics, and industrial\nrobotics. Failures of such systems can lead to loss of life or property, which\nnecessitates stringent verification and validation for providing high\nassurance. Though formal verification approaches are being investigated,\ntesting remains the primary technique for assessing the dependability of such\nsystems. Due to the nature of the tasks handled by DNNs, the cost of obtaining\ntest oracle data---the expected output, a.k.a. label, for a given input---is\nhigh, which significantly impacts the amount and quality of testing that can be\nperformed. Thus, prioritizing input data for testing DNNs in meaningful ways to\nreduce the cost of labeling can go a long way in increasing testing efficacy.\nThis paper proposes using gauges of the DNN's sentiment derived from the\ncomputation performed by the model, as a means to identify inputs that are\nlikely to reveal weaknesses. We empirically assessed the efficacy of three such\nsentiment measures for prioritization---confidence, uncertainty, and\nsurprise---and compare their effectiveness in terms of their fault-revealing\ncapability and retraining effectiveness. The results indicate that sentiment\nmeasures can effectively flag inputs that expose unacceptable DNN behavior. For\nMNIST models, the average percentage of inputs correctly flagged ranged from\n88% to 94.8%.","url_abs":"http://arxiv.org/abs/1901.03768v1","url_pdf":"http://arxiv.org/pdf/1901.03768v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"input-prioritization-for-testing-neural","repo_url":"https://github.com/bntejn/keras-prioritizer","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"self-driving-cars","task_name":"Self-Driving Cars"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1901.03768","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}