{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-deep-recurrent-q-network-towards-self","title":"A Deep Recurrent Q Network towards Self-adapting Distributed Microservices architecture","arxiv_id":"1901.04011","date":"2019-01-13","proceeding":null,"authors":["Basel Magableh"],"abstract":"One desired aspect of microservices architecture is the ability to self-adapt\nits own architecture and behaviour in response to changes in the operational\nenvironment. To achieve the desired high levels of self-adaptability, this\nresearch implements the distributed microservices architectures model, as\ninformed by the MAPE-K model. The proposed architecture employs a multi\nadaptation agents supported by a centralised controller, that can observe the\nenvironment and execute a suitable adaptation action. The adaptation planning\nis managed by a deep recurrent Q-network (DRQN). It is argued that such\nintegration between DRQN and MDP agents in a MAPE-K model offers distributed\nmicroservice architecture with self-adaptability and high levels of\navailability and scalability. Integrating DRQN into the adaptation process\nimproves the effectiveness of the adaptation and reduces any adaptation risks,\nincluding resources over-provisioning and thrashing. The performance of DRQN is\nevaluated against deep Q-learning and policy gradient algorithms including: i)\ndeep q-network (DQN), ii) dulling deep Q-network (DDQN), iii) a policy gradient\nneural network (PGNN), and iv) deep deterministic policy gradient (DDPG). The\nDRQN implementation in this paper manages to outperform the above mentioned\nalgorithms in terms of total reward, less adaptation time, lower error rates,\nplus faster convergence and training times. We strongly believe that DRQN is\nmore suitable for driving the adaptation in distributed services-oriented\narchitecture and offers better performance than other dynamic decision-making\nalgorithms.","url_abs":"http://arxiv.org/abs/1901.04011v2","url_pdf":"http://arxiv.org/pdf/1901.04011v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-deep-recurrent-q-network-towards-self","repo_url":"https://github.com/baselm/cmarl","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"q-learning","task_name":"Q-Learning"}],"methods":[{"method_slug":"q-learning","method_name":"Q-Learning"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}