{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/decentralized-computation-offloading-for","title":"Decentralized Computation Offloading for Multi-User Mobile Edge Computing: A Deep Reinforcement Learning Approach","arxiv_id":"1812.07394","date":"2018-12-16","proceeding":null,"authors":["Zhao Chen","Xiaodong Wang"],"abstract":"Mobile edge computing (MEC) emerges recently as a promising solution to\nrelieve resource-limited mobile devices from computation-intensive tasks, which\nenables devices to offload workloads to nearby MEC servers and improve the\nquality of computation experience. Nevertheless, by considering a MEC system\nconsisting of multiple mobile users with stochastic task arrivals and wireless\nchannels in this paper, the design of computation offloading policies is\nchallenging to minimize the long-term average computation cost in terms of\npower consumption and buffering delay. A deep reinforcement learning (DRL)\nbased decentralized dynamic computation offloading strategy is investigated to\nbuild a scalable MEC system with limited feedback. Specifically, a continuous\naction space-based DRL approach named deep deterministic policy gradient (DDPG)\nis adopted to learn efficient computation offloading policies independently at\neach mobile user. Thus, powers of both local execution and task offloading can\nbe adaptively allocated by the learned policies from each user's local\nobservation of the MEC system. Numerical results are illustrated to demonstrate\nthat efficient policies can be learned at each user, and performance of the\nproposed DDPG based decentralized strategy outperforms the conventional deep\nQ-network (DQN) based discrete power control strategy and some other greedy\nstrategies with reduced computation cost. Besides, the power-delay tradeoff is\nalso analyzed for both the DDPG based and DQN based strategies.","url_abs":"http://arxiv.org/abs/1812.07394v1","url_pdf":"http://arxiv.org/pdf/1812.07394v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"decentralized-computation-offloading-for","repo_url":"https://github.com/Vatsala17/Vatsala17","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"decentralized-computation-offloading-for","repo_url":"https://github.com/swordest/mec_drl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"edge-computing","task_name":"Edge-computing"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"}],"methods":[{"method_slug":"adam","method_name":"Adam"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"ddpg","method_name":"DDPG"},{"method_slug":"dqn","method_name":"DQN"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"experience-replay","method_name":"Experience Replay"},{"method_slug":"q-learning","method_name":"Q-Learning"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"weight-decay","method_name":"Weight Decay"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}