{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/differentiable-fine-grained-quantization-for","title":"Differentiable Fine-grained Quantization for Deep Neural Network Compression","arxiv_id":"1810.10351","date":"2018-10-20","proceeding":"NIPS Workshop CDNNRIA 2018","authors":["Hsin-Pai Cheng","Yuanjun Huang","Xuyang Guo","Yifei HUANG","Feng Yan","Hai Li","Yiran Chen"],"abstract":"Neural networks have shown great performance in cognitive tasks. When\ndeploying network models on mobile devices with limited resources, weight\nquantization has been widely adopted. Binary quantization obtains the highest\ncompression but usually results in big accuracy drop. In practice, 8-bit or\n16-bit quantization is often used aiming at maintaining the same accuracy as\nthe original 32-bit precision. We observe different layers have different\naccuracy sensitivity of quantization. Thus judiciously selecting different\nprecision for different layers/structures can potentially produce more\nefficient models compared to traditional quantization methods by striking a\nbetter balance between accuracy and compression rate. In this work, we propose\na fine-grained quantization approach for deep neural network compression by\nrelaxing the search space of quantization bitwidth from discrete to a\ncontinuous domain. The proposed approach applies gradient descend based\noptimization to generate a mixed-precision quantization scheme that outperforms\nthe accuracy of traditional quantization methods under the same compression\nrate.","url_abs":"http://arxiv.org/abs/1810.10351v3","url_pdf":"http://arxiv.org/pdf/1810.10351v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"differentiable-fine-grained-quantization-for","repo_url":"https://github.com/newwhitecheng/compress-all-nn","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"neural-network-compression","task_name":"Neural Network Compression"},{"task_slug":"quantization","task_name":"Quantization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1810.10351","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}