{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-model-based-strategies-in-simple","title":"Learning model-based strategies in simple environments with hierarchical q-networks","arxiv_id":"1801.06689","date":"2018-01-20","proceeding":null,"authors":["Necati Alp Muyesser","Kyle Dunovan","Timothy Verstynen"],"abstract":"Recent advances in deep learning have allowed artificial agents to rival\nhuman-level performance on a wide range of complex tasks; however, the ability\nof these networks to learn generalizable strategies remains a pressing\nchallenge. This critical limitation is due in part to two factors: the opaque\ninformation representation in deep neural networks and the complexity of the\ntask environments in which they are typically deployed. Here we propose a novel\nHierarchical Q-Network (HQN) motivated by theories of the hierarchical\norganization of the human prefrontal cortex, that attempts to identify lower\ndimensional patterns in the value landscape that can be exploited to construct\nan internal model of rules in simple environments. We draw on combinatorial\ngames, where there exists a single optimal strategy for winning that\ngeneralizes across other features of the game, to probe the strategy\ngeneralization of the HQN and other reinforcement learning (RL) agents using\nvariations of Wythoff's game. Traditional RL approaches failed to reach\nsatisfactory performance on variants of Wythoff's Game; however, the HQN\nlearned heuristic-like strategies that generalized across changes in board\nconfiguration. More importantly, the HQN allowed for transparent inspection of\nthe agent's internal model of the game following training. Our results show how\na biologically inspired hierarchical learner can facilitate learning abstract\nrules to promote robust and flexible action policies in simplified training\nenvironments with clearly delineated optimal strategies.","url_abs":"http://arxiv.org/abs/1801.06689v1","url_pdf":"http://arxiv.org/pdf/1801.06689v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-model-based-strategies-in-simple","repo_url":"https://github.com/CoAxLab/azad","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}