{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/variational-information-maximisation-for","title":"Variational Information Maximisation for Intrinsically Motivated Reinforcement Learning","arxiv_id":"1509.08731","date":"2015-09-29","proceeding":"NeurIPS 2015 12","authors":["Shakir Mohamed","Danilo Jimenez Rezende"],"abstract":"The mutual information is a core statistical quantity that has applications\nin all areas of machine learning, whether this is in training of density models\nover multiple data modalities, in maximising the efficiency of noisy\ntransmission channels, or when learning behaviour policies for exploration by\nartificial agents. Most learning algorithms that involve optimisation of the\nmutual information rely on the Blahut-Arimoto algorithm --- an enumerative\nalgorithm with exponential complexity that is not suitable for modern machine\nlearning applications. This paper provides a new approach for scalable\noptimisation of the mutual information by merging techniques from variational\ninference and deep learning. We develop our approach by focusing on the problem\nof intrinsically-motivated learning, where the mutual information forms the\ndefinition of a well-known internal drive known as empowerment. Using a\nvariational lower bound on the mutual information, combined with convolutional\nnetworks for handling visual input streams, we develop a stochastic\noptimisation algorithm that allows for scalable information maximisation and\nempowerment-based reasoning directly from pixels to actions.","url_abs":"http://arxiv.org/abs/1509.08731v1","url_pdf":"http://arxiv.org/pdf/1509.08731v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"variational-information-maximisation-for","repo_url":"https://github.com/AidanRocke/variational_empowerment","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"variational-information-maximisation-for","repo_url":"https://github.com/Lham71/Maximising-Neural-Information-Processing","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"variational-inference","task_name":"Variational Inference"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1509.08731","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}