{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/maxdnn-an-efficient-convolution-kernel-for","title":"maxDNN: An Efficient Convolution Kernel for Deep Learning with Maxwell GPUs","arxiv_id":"1501.06633","date":"2015-01-27","proceeding":null,"authors":["Andrew Lavin"],"abstract":"This paper describes maxDNN, a computationally efficient convolution kernel\nfor deep learning with the NVIDIA Maxwell GPU. maxDNN reaches 96.3%\ncomputational efficiency on typical deep learning network architectures. The\ndesign combines ideas from cuda-convnet2 with the Maxas SGEMM assembly code. We\nonly address forward propagation (FPROP) operation of the network, but we\nbelieve that the same techniques used here will be effective for backward\npropagation (BPROP) as well.","url_abs":"http://arxiv.org/abs/1501.06633v3","url_pdf":"http://arxiv.org/pdf/1501.06633v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"maxdnn-an-efficient-convolution-kernel-for","repo_url":"https://github.com/eBay/maxDNN","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"computational-efficiency","task_name":"Computational Efficiency"},{"task_slug":"deep-learning","task_name":"Deep Learning"},{"task_slug":null,"task_name":"GPU"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}