{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-reinforcement-learning-for-cybersecurity","title":"Deep Reinforcement Learning for Cybersecurity Assessment of Wind Integrated Power Systems","arxiv_id":"2007.03025","date":"2020-11-15","proceeding":null,"authors":[],"abstract":"The integration of renewable energy sources (RES) is rapidly increasing in\nelectric power systems (EPS). While the inclusion of intermittent RES coupled\nwith the wide-scale deployment of communication and sensing devices is\nimportant towards a fully smart grid, it has also expanded the cyber-threat\nlandscape, effectively making power systems vulnerable to cyberattacks. This\npaper proposes a cybersecurity assessment approach designed to assess the\ncyberphysical security of EPS. The work takes into consideration the\nintermittent generation of RES, vulnerabilities introduced by\nmicroprocessor-based electronic information and operational technology (IT/OT)\ndevices, and contingency analysis results. The proposed approach utilizes deep\nreinforcement learning (DRL) and an adapted Common Vulnerability Scoring System\n(CVSS) score tailored to assess vulnerabilities in EPS in order to identify the\noptimal attack transition policy based on N-2 contingency results, i.e., the\nsimultaneous failure of two system elements. The effectiveness of the work is\nvalidated via numerical and real-time simulation experiments performed on\nliterature-based power grid test cases. The results demonstrate how the\nproposed method based on deep Q-network (DQN) performs closely to a\ngraph-search approach in terms of the number of transitions needed to find the\noptimal attack policy, without the need for full observation of the system. In\naddition, the experiments present the method's scalability by showcasing the\nnumber of transitions needed to find the optimal attack transition policy in a\nlarge system such as the Polish 2383 bus test system. The results exhibit how\nthe proposed approach requires one order of magnitude fewer transitions when\ncompared to a random transition policy.","url_abs":"http://arxiv.org/abs/2007.03025v2","url_pdf":"http://arxiv.org/pdf/2007.03025v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-reinforcement-learning-for-cybersecurity","repo_url":"https://github.com/DSS-lab/DRLCyberAssessment_DQNCode","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}