{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/on-loss-functions-for-deep-neural-networks-in","title":"On Loss Functions for Deep Neural Networks in Classification","arxiv_id":"1702.05659","date":"2017-02-18","proceeding":null,"authors":["Katarzyna Janocha","Wojciech Marian Czarnecki"],"abstract":"Deep neural networks are currently among the most commonly used classifiers.\nDespite easily achieving very good performance, one of the best selling points\nof these models is their modular design - one can conveniently adapt their\narchitecture to specific needs, change connectivity patterns, attach\nspecialised layers, experiment with a large amount of activation functions,\nnormalisation schemes and many others. While one can find impressively wide\nspread of various configurations of almost every aspect of the deep nets, one\nelement is, in authors' opinion, underrepresented - while solving\nclassification problems, vast majority of papers and applications simply use\nlog loss. In this paper we try to investigate how particular choices of loss\nfunctions affect deep models and their learning dynamics, as well as resulting\nclassifiers robustness to various effects. We perform experiments on classical\ndatasets, as well as provide some additional, theoretical insights into the\nproblem. In particular we show that L1 and L2 losses are, quite surprisingly,\njustified classification objectives for deep nets, by providing probabilistic\ninterpretation in terms of expected misclassification. We also introduce two\nlosses which are not typically used as deep nets objectives and show that they\nare viable alternatives to the existing ones.","url_abs":"http://arxiv.org/abs/1702.05659v1","url_pdf":"http://arxiv.org/pdf/1702.05659v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"on-loss-functions-for-deep-neural-networks-in","repo_url":"https://github.com/raj2022/Basics-","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1702.05659","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}