{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/emulating-malware-authors-for-proactive","title":"Emulating malware authors for proactive protection using GANs over a distributed image visualization of dynamic file behavior","arxiv_id":"1807.07525","date":"2018-07-19","proceeding":null,"authors":["Vineeth S. Bhaskara","Debanjan Bhattacharyya"],"abstract":"Malware authors have always been at an advantage of being able to\nadversarially test and augment their malicious code, before deploying the\npayload, using anti-malware products at their disposal. The anti-malware\ndevelopers and threat experts, on the other hand, do not have such a privilege\nof tuning anti-malware products against zero-day attacks pro-actively. This\nallows the malware authors to being a step ahead of the anti-malware products,\nfundamentally biasing the cat and mouse game played by the two parties. In this\npaper, we propose a way that would enable machine learning based threat\nprevention models to bridge that gap by being able to tune against a deep\ngenerative adversarial network (GAN), which takes up the role of a malware\nauthor and generates new types of malware. The GAN is trained over a reversible\ndistributed RGB image representation of known malware behaviors, encoding the\nsequence of API call ngrams and the corresponding term frequencies. The\ngenerated images represent synthetic malware that can be decoded back to the\nunderlying API call sequence information. The image representation is not only\ndemonstrated as a general technique of incorporating necessary priors for\nexploiting convolutional neural network architectures for generative or\ndiscriminative modeling, but also as a visualization method for easy manual\nsoftware or malware categorization, by having individual API ngram information\ndistributed across the image space. In addition, we also propose using\nsmart-definitions for detecting malwares based on perceptual hashing of these\nimages. Such hashes are potentially more effective than cryptographic hashes\nthat do not carry any meaningful similarity metric, and hence, do not\ngeneralize well.","url_abs":"http://arxiv.org/abs/1807.07525v2","url_pdf":"http://arxiv.org/pdf/1807.07525v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"emulating-malware-authors-for-proactive","repo_url":"https://github.com/bsvineethiitg/malwaregan","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":null,"task_name":"Generative Adversarial Network"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}