{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/utilizing-class-information-for-deep-network","title":"Utilizing Class Information for Deep Network Representation Shaping","arxiv_id":"1809.09307","date":"2018-09-25","proceeding":null,"authors":["Daeyoung Choi","Wonjong Rhee"],"abstract":"Statistical characteristics of deep network representations, such as sparsity\nand correlation, are known to be relevant to the performance and\ninterpretability of deep learning. When a statistical characteristic is\ndesired, often an adequate regularizer can be designed and applied during the\ntraining phase. Typically, such a regularizer aims to manipulate a statistical\ncharacteristic over all classes together. For classification tasks, however, it\nmight be advantageous to enforce the desired characteristic per class such that\ndifferent classes can be better distinguished. Motivated by the idea, we design\ntwo class-wise regularizers that explicitly utilize class information:\nclass-wise Covariance Regularizer (cw-CR) and class-wise Variance Regularizer\n(cw-VR). cw-CR targets to reduce the covariance of representations calculated\nfrom the same class samples for encouraging feature independence. cw-VR is\nsimilar, but variance instead of covariance is targeted to improve feature\ncompactness. For the sake of completeness, their counterparts without using\nclass information, Covariance Regularizer (CR) and Variance Regularizer (VR),\nare considered together. The four regularizers are conceptually simple and\ncomputationally very efficient, and the visualization shows that the\nregularizers indeed perform distinct representation shaping. In terms of\nclassification performance, significant improvements over the baseline and\nL1/L2 weight regularization methods were found for 21 out of 22 tasks over\npopular benchmark datasets. In particular, cw-VR achieved the best performance\nfor 13 tasks including ResNet-32/110.","url_abs":"http://arxiv.org/abs/1809.09307v2","url_pdf":"http://arxiv.org/pdf/1809.09307v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"utilizing-class-information-for-deep-network","repo_url":"https://github.com/snu-adsl/class_wise_regularizer","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"classification","task_name":"General Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}