{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/recurrent-neural-network-language-models-for","title":"Recurrent Neural Network Language Models for Open Vocabulary Event-Level Cyber Anomaly Detection","arxiv_id":"1712.00557","date":"2017-12-02","proceeding":null,"authors":["Aaron Tuor","Ryan Baerwolf","Nicolas Knowles","Brian Hutchinson","Nicole Nichols","Rob Jasper"],"abstract":"Automated analysis methods are crucial aids for monitoring and defending a\nnetwork to protect the sensitive or confidential data it hosts. This work\nintroduces a flexible, powerful, and unsupervised approach to detecting\nanomalous behavior in computer and network logs, one that largely eliminates\ndomain-dependent feature engineering employed by existing methods. By treating\nsystem logs as threads of interleaved \"sentences\" (event log lines) to train\nonline unsupervised neural network language models, our approach provides an\nadaptive model of normal network behavior. We compare the effectiveness of both\nstandard and bidirectional recurrent neural network language models at\ndetecting malicious activity within network log data. Extending these models,\nwe introduce a tiered recurrent architecture, which provides context by\nmodeling sequences of users' actions over time. Compared to Isolation Forest\nand Principal Components Analysis, two popular anomaly detection algorithms, we\nobserve superior performance on the Los Alamos National Laboratory Cyber\nSecurity dataset. For log-line-level red team detection, our best performing\ncharacter-based model provides test set area under the receiver operator\ncharacteristic curve of 0.98, demonstrating the strong fine-grained anomaly\ndetection performance of this approach on open vocabulary logging sources.","url_abs":"http://arxiv.org/abs/1712.00557v1","url_pdf":"http://arxiv.org/pdf/1712.00557v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"recurrent-neural-network-language-models-for","repo_url":"https://github.com/pnnl/safekit","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"anomaly-detection","task_name":"Anomaly Detection"},{"task_slug":"feature-engineering","task_name":"Feature Engineering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1712.00557","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}