{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/robust-neural-malware-detection-models-for","title":"Robust Neural Malware Detection Models for Emulation Sequence Learning","arxiv_id":"1806.10741","date":"2018-06-28","proceeding":null,"authors":["Rakshit Agrawal","Jack W. Stokes","Mady Marinescu","Karthik Selvaraj"],"abstract":"Malicious software, or malware, presents a continuously evolving challenge in\ncomputer security. These embedded snippets of code in the form of malicious\nfiles or hidden within legitimate files cause a major risk to systems with\ntheir ability to run malicious command sequences. Malware authors even use\npolymorphism to reorder these commands and create several malicious variations.\nHowever, if executed in a secure environment, one can perform early malware\ndetection on emulated command sequences.\n  The models presented in this paper leverage this sequential data derived via\nemulation in order to perform Neural Malware Detection. These models target the\ncore of the malicious operation by learning the presence and pattern of\nco-occurrence of malicious event actions from within these sequences. Our\nmodels can capture entire event sequences and be trained directly using the\nknown target labels. These end-to-end learning models are powered by two\ncommonly used structures - Long Short-Term Memory (LSTM) Networks and\nConvolutional Neural Networks (CNNs). Previously proposed sequential malware\nclassification models process no more than 200 events. Attackers can evade\ndetection by delaying any malicious activity beyond the beginning of the file.\nWe present specialized models that can handle extremely long sequences while\nsuccessfully performing malware detection in an efficient way. We present an\nimplementation of the Convoluted Partitioning of Long Sequences approach in\norder to tackle this vulnerability and operate on long sequences. We present\nour results on a large dataset consisting of 634,249 file sequences, with\nextremely long file sequences.","url_abs":"http://arxiv.org/abs/1806.10741v1","url_pdf":"http://arxiv.org/pdf/1806.10741v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"robust-neural-malware-detection-models-for","repo_url":"https://github.com/tychen5/sportslottery","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"computer-security","task_name":"Computer Security"},{"task_slug":"malware-classification","task_name":"Malware Classification"},{"task_slug":"malware-detection","task_name":"Malware Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}