{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/highly-efficient-8-bit-low-precision-1","title":"Highly Efficient 8-bit Low Precision Inference of Convolutional Neural Networks with IntelCaffe","arxiv_id":"1805.08691","date":"2018-05-04","proceeding":null,"authors":["Jiong Gong","Haihao Shen","Guoming Zhang","Xiaoli Liu","Shane Li","Ge Jin","Niharika Maheshwari","Evarist Fomenko","Eden Segal"],"abstract":"High throughput and low latency inference of deep neural networks are\ncritical for the deployment of deep learning applications. This paper presents\nthe efficient inference techniques of IntelCaffe, the first Intel optimized\ndeep learning framework that supports efficient 8-bit low precision inference\nand model optimization techniques of convolutional neural networks on Intel\nXeon Scalable Processors. The 8-bit optimized model is automatically generated\nwith a calibration process from FP32 model without the need of fine-tuning or\nretraining. We show that the inference throughput and latency with ResNet-50,\nInception-v3 and SSD are improved by 1.38X-2.9X and 1.35X-3X respectively with\nneglectable accuracy loss from IntelCaffe FP32 baseline and by 56X-75X and\n26X-37X from BVLC Caffe. All these techniques have been open-sourced on\nIntelCaffe GitHub1, and the artifact is provided to reproduce the result on\nAmazon AWS Cloud.","url_abs":"http://arxiv.org/abs/1805.08691v1","url_pdf":"http://arxiv.org/pdf/1805.08691v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"highly-efficient-8-bit-low-precision-1","repo_url":"https://github.com/intel/caffe","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"deep-learning","task_name":"Deep Learning"},{"task_slug":"model-optimization","task_name":"Model Optimization"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"non-maximum-suppression","method_name":"Non Maximum Suppression"},{"method_slug":"ssd","method_name":"SSD"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}