{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/semi-supervised-deep-reinforcement-learning","title":"Semi-supervised Deep Reinforcement Learning in Support of IoT and Smart City Services","arxiv_id":"1810.04118","date":"2018-10-09","proceeding":null,"authors":["Mehdi Mohammadi","Ala Al-Fuqaha","Mohsen Guizani","Jun-Seok Oh"],"abstract":"Smart services are an important element of the smart cities and the Internet\nof Things (IoT) ecosystems where the intelligence behind the services is\nobtained and improved through the sensory data. Providing a large amount of\ntraining data is not always feasible; therefore, we need to consider\nalternative ways that incorporate unlabeled data as well. In recent years, Deep\nreinforcement learning (DRL) has gained great success in several application\ndomains. It is an applicable method for IoT and smart city scenarios where\nauto-generated data can be partially labeled by users' feedback for training\npurposes. In this paper, we propose a semi-supervised deep reinforcement\nlearning model that fits smart city applications as it consumes both labeled\nand unlabeled data to improve the performance and accuracy of the learning\nagent. The model utilizes Variational Autoencoders (VAE) as the inference\nengine for generalizing optimal policies. To the best of our knowledge, the\nproposed model is the first investigation that extends deep reinforcement\nlearning to the semi-supervised paradigm. As a case study of smart city\napplications, we focus on smart buildings and apply the proposed model to the\nproblem of indoor localization based on BLE signal strength. Indoor\nlocalization is the main component of smart city services since people spend\nsignificant time in indoor environments. Our model learns the best action\npolicies that lead to a close estimation of the target locations with an\nimprovement of 23% in terms of distance to the target and at least 67% more\nreceived rewards compared to the supervised DRL model.","url_abs":"http://arxiv.org/abs/1810.04118v1","url_pdf":"http://arxiv.org/pdf/1810.04118v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"semi-supervised-deep-reinforcement-learning","repo_url":"https://github.com/nikola310/indoor-localization","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"indoor-localization","task_name":"Indoor Localization"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}