{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/adaptive-coordination-of-working-memory-and","title":"Adaptive coordination of working-memory and reinforcement learning in non-human primates performing a trial-and-error problem solving task","arxiv_id":"1711.00698","date":"2017-11-02","proceeding":null,"authors":["Guillaume Viejo","Benoît Girard","Emmanuel Procyk","Mehdi Khamassi"],"abstract":"Accumulating evidence suggest that human behavior in trial-and-error learning\ntasks based on decisions between discrete actions may involve a combination of\nreinforcement learning (RL) and working-memory (WM). While the understanding of\nbrain activity at stake in this type of tasks often involve the comparison with\nnon-human primate neurophysiological results, it is not clear whether monkeys\nuse similar combined RL and WM processes to solve these tasks. Here we analyzed\nthe behavior of five monkeys with computational models combining RL and WM. Our\nmodel-based analysis approach enables to not only fit trial-by-trial choices\nbut also transient slowdowns in reaction times, indicative of WM use. We found\nthat the behavior of the five monkeys was better explained in terms of a\ncombination of RL and WM despite inter-individual differences. The same\ncoordination dynamics we used in a previous study in humans best explained the\nbehavior of some monkeys while the behavior of others showed the opposite\npattern, revealing a possible different dynamics of WM process. We further\nanalyzed different variants of the tested models to open a discussion on how\nthe long pretraining in these tasks may have favored particular coordination\ndynamics between RL and WM. This points towards either inter-species\ndifferences or protocol differences which could be further tested in humans.","url_abs":"http://arxiv.org/abs/1711.00698v1","url_pdf":"http://arxiv.org/pdf/1711.00698v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"adaptive-coordination-of-working-memory-and","repo_url":"https://github.com/gviejo/Gohal","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"}],"methods":[{"method_slug":"q-learning","method_name":"Q-Learning"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}