{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/improving-the-neural-gpu-architecture-for","title":"Improving the Neural GPU Architecture for Algorithm Learning","arxiv_id":"1702.08727","date":"2017-02-28","proceeding":null,"authors":["Karlis Freivalds","Renars Liepins"],"abstract":"Algorithm learning is a core problem in artificial intelligence with\nsignificant implications on automation level that can be achieved by machines.\nRecently deep learning methods are emerging for synthesizing an algorithm from\nits input-output examples, the most successful being the Neural GPU, capable of\nlearning multiplication. We present several improvements to the Neural GPU that\nsubstantially reduces training time and improves generalization. We introduce a\nnew technique - hard nonlinearities with saturation costs- that has general\napplicability. We also introduce a technique of diagonal gates that can be\napplied to active-memory models. The proposed architecture is the first capable\nof learning decimal multiplication end-to-end.","url_abs":"http://arxiv.org/abs/1702.08727v2","url_pdf":"http://arxiv.org/pdf/1702.08727v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"improving-the-neural-gpu-architecture-for","repo_url":"https://github.com/LUMII-Syslab/DNGPU","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"improving-the-neural-gpu-architecture-for","repo_url":"https://github.com/LUMII-Syslab/shuffle-exchange","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":null,"task_name":"GPU"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}