{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/enabling-sparse-winograd-convolution-by","title":"Enabling Sparse Winograd Convolution by Native Pruning","arxiv_id":"1702.08597","date":"2017-02-28","proceeding":null,"authors":["Sheng Li","Jongsoo Park","Ping Tak Peter Tang"],"abstract":"Sparse methods and the use of Winograd convolutions are two orthogonal\napproaches, each of which significantly accelerates convolution computations in\nmodern CNNs. Sparse Winograd merges these two and thus has the potential to\noffer a combined performance benefit. Nevertheless, training convolution layers\nso that the resulting Winograd kernels are sparse has not hitherto been very\nsuccessful. By introducing a Winograd layer in place of a standard convolution\nlayer, we can learn and prune Winograd coefficients \"natively\" and obtain\nsparsity level beyond 90% with only 0.1% accuracy loss with AlexNet on ImageNet\ndataset. Furthermore, we present a sparse Winograd convolution algorithm and\nimplementation that exploits the sparsity, achieving up to 31.7 effective\nTFLOP/s in 32-bit precision on a latest Intel Xeon CPU, which corresponds to a\n5.4x speedup over a state-of-the-art dense convolution implementation.","url_abs":"http://arxiv.org/abs/1702.08597v2","url_pdf":"http://arxiv.org/pdf/1702.08597v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"enabling-sparse-winograd-convolution-by","repo_url":"https://github.com/IntelLabs/SkimCaffe","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"gone","observed_at":"2026-09-18","how":"tree_404+repo_404"}}],"tasks":[{"task_slug":null,"task_name":"CPU"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"grouped-convolution","method_name":"Grouped Convolution"},{"method_slug":"local-response-normalization","method_name":"Local Response Normalization"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1702.08597","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}