{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/information-theoretic-understanding-of","title":"Information-Theoretic Understanding of Population Risk Improvement with Model Compression","arxiv_id":"1901.09421","date":"2019-01-27","proceeding":null,"authors":["Yuheng Bu","Weihao Gao","Shaofeng Zou","Venugopal V. Veeravalli"],"abstract":"We show that model compression can improve the population risk of a\npre-trained model, by studying the tradeoff between the decrease in the\ngeneralization error and the increase in the empirical risk with model\ncompression. We first prove that model compression reduces an\ninformation-theoretic bound on the generalization error; this allows for an\ninterpretation of model compression as a regularization technique to avoid\noverfitting. We then characterize the increase in empirical risk with model\ncompression using rate distortion theory. These results imply that the\npopulation risk could be improved by model compression if the decrease in\ngeneralization error exceeds the increase in empirical risk. We show through a\nlinear regression example that such a decrease in population risk due to model\ncompression is indeed possible. Our theoretical results further suggest that\nthe Hessian-weighted $K$-means clustering compression approach can be improved\nby regularizing the distance between the clustering centers. We provide\nexperiments with neural networks to support our theoretical assertions.","url_abs":"http://arxiv.org/abs/1901.09421v1","url_pdf":"http://arxiv.org/pdf/1901.09421v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"information-theoretic-understanding-of","repo_url":"https://github.com/wgao9/weight_quant","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"model-compression","task_name":"Model Compression"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1901.09421","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}