{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/scaling-memory-augmented-neural-networks-with","title":"Scaling Memory-Augmented Neural Networks with Sparse Reads and Writes","arxiv_id":"1610.09027","date":"2016-10-27","proceeding":"NeurIPS 2016 12","authors":["Jack W. Rae","Jonathan J. Hunt","Tim Harley","Ivo Danihelka","Andrew Senior","Greg Wayne","Alex Graves","Timothy P. Lillicrap"],"abstract":"Neural networks augmented with external memory have the ability to learn\nalgorithmic solutions to complex tasks. These models appear promising for\napplications such as language modeling and machine translation. However, they\nscale poorly in both space and time as the amount of memory grows --- limiting\ntheir applicability to real-world domains. Here, we present an end-to-end\ndifferentiable memory access scheme, which we call Sparse Access Memory (SAM),\nthat retains the representational power of the original approaches whilst\ntraining efficiently with very large memories. We show that SAM achieves\nasymptotic lower bounds in space and time complexity, and find that an\nimplementation runs $1,\\!000\\times$ faster and with $3,\\!000\\times$ less\nphysical memory than non-sparse models. SAM learns with comparable data\nefficiency to existing models on a range of synthetic tasks and one-shot\nOmniglot character recognition, and can scale to tasks requiring $100,\\!000$s\nof time steps and memories. As well, we show how our approach can be adapted\nfor models that maintain temporal associations between memories, as with the\nrecently introduced Differentiable Neural Computer.","url_abs":"http://arxiv.org/abs/1610.09027v1","url_pdf":"http://arxiv.org/pdf/1610.09027v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"translation","task_name":"Translation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/question-answering-on-babi","task":"Question Answering","dataset":"bAbi","model":"LSTM","rank_in_archive_order":10,"of":14,"metrics":{"Accuracy (trained on 1k)":"49%","Mean Error Rate":"28.7%"},"uses_additional_data":false},{"leaderboard":"/sota/question-answering-on-babi","task":"Question Answering","dataset":"bAbi","model":"SDNC","rank_in_archive_order":14,"of":14,"metrics":{"Mean Error Rate":"6.4%"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1610.09027","atlas_url":"https://app.syntology.ai/?focus=1610.09027","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}