{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/trojdrl-trojan-attacks-on-deep-reinforcement","title":"TrojDRL: Trojan Attacks on Deep Reinforcement Learning Agents","arxiv_id":"1903.06638","date":"2019-03-01","proceeding":null,"authors":["Panagiota Kiourti","Kacper Wardega","Susmit Jha","Wenchao Li"],"abstract":"Recent work has identified that classification models implemented as neural\nnetworks are vulnerable to data-poisoning and Trojan attacks at training time.\nIn this work, we show that these training-time vulnerabilities extend to deep\nreinforcement learning (DRL) agents and can be exploited by an adversary with\naccess to the training process. In particular, we focus on Trojan attacks that\naugment the function of reinforcement learning policies with hidden behaviors.\nWe demonstrate that such attacks can be implemented through minuscule data\npoisoning (as little as 0.025% of the training data) and in-band reward\nmodification that does not affect the reward on normal inputs. The policies\nlearned with our proposed attack approach perform imperceptibly similar to\nbenign policies but deteriorate drastically when the Trojan is triggered in\nboth targeted and untargeted settings. Furthermore, we show that existing\nTrojan defense mechanisms for classification tasks are not effective in the\nreinforcement learning setting.","url_abs":"http://arxiv.org/abs/1903.06638v1","url_pdf":"http://arxiv.org/pdf/1903.06638v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"trojdrl-trojan-attacks-on-deep-reinforcement","repo_url":"https://github.com/pkiourti/rl_backdoor","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"trojdrl-trojan-attacks-on-deep-reinforcement","repo_url":"https://github.com/nwpuhkp/DRL-Backdoor-Pong","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"data-poisoning","task_name":"Data Poisoning"},{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1903.06638","atlas_url":"https://app.syntology.ai/?focus=1903.06638","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}