{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cyberseceval-2-a-wide-ranging-cybersecurity","title":"CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models","arxiv_id":"2404.13161","date":"2024-04-19","proceeding":null,"authors":["Manish Bhatt","Sahana Chennabasappa","Yue Li","Cyrus Nikolaidis","Daniel Song","Shengye Wan","Faizan Ahmad","Cornelius Aschermann","Yaohui Chen","Dhaval Kapil","David Molnar","Spencer Whitman","Joshua Saxe"],"abstract":"Large language models (LLMs) introduce new security risks, but there are few comprehensive evaluation suites to measure and reduce these risks. We present BenchmarkName, a novel benchmark to quantify LLM security risks and capabilities. We introduce two new areas for testing: prompt injection and code interpreter abuse. We evaluated multiple state-of-the-art (SOTA) LLMs, including GPT-4, Mistral, Meta Llama 3 70B-Instruct, and Code Llama. Our results show that conditioning away risk of attack remains an unsolved problem; for example, all tested models showed between 26% and 41% successful prompt injection tests. We further introduce the safety-utility tradeoff: conditioning an LLM to reject unsafe prompts can cause the LLM to falsely reject answering benign prompts, which lowers utility. We propose quantifying this tradeoff using False Refusal Rate (FRR). As an illustration, we introduce a novel test set to quantify FRR for cyberattack helpfulness risk. We find many LLMs able to successfully comply with \"borderline\" benign requests while still rejecting most unsafe requests. Finally, we quantify the utility of LLMs for automating a core cybersecurity task, that of exploiting software vulnerabilities. This is important because the offensive capabilities of LLMs are of intense interest; we quantify this by creating novel test sets for four representative problems. We find that models with coding capabilities perform better than those without, but that further work is needed for LLMs to become proficient at exploit generation. Our code is open source and can be used to evaluate other LLMs.","url_abs":"https://arxiv.org/abs/2404.13161v1","url_pdf":"https://arxiv.org/pdf/2404.13161v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"cyberseceval-2-a-wide-ranging-cybersecurity","repo_url":"https://github.com/facebookresearch/purplellama","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"gpt-4","method_name":"GPT-4"},{"method_slug":"llama","method_name":"LLaMA"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"set","method_name":"SET"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2404.13161","atlas_url":"https://app.syntology.ai/?focus=2404.13161","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2404.13161"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/facebookresearch/purplellama","reach":{"status":"ok","spdx":"NOASSERTION"}}],"summary":{"ran":3,"unverified":1},"by_repo_kind":{"official":{"samples":4,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"67b1c1f6d748c720","entry":"filter_frame","repo":"facebookresearch/purplellama","repo_kind":"official","path":"CybersecurityBenchmarks/benchmark/arvo_utils.py","file_url":"https://github.com/facebookresearch/purplellama/blob/HEAD/CybersecurityBenchmarks/benchmark/arvo_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"67b1c1f6d748c720"}},{"code_sha256_prefix":"744a6e2320d37acf","entry":"filter_traces","repo":"facebookresearch/purplellama","repo_kind":"official","path":"CybersecurityBenchmarks/benchmark/arvo_utils.py","file_url":"https://github.com/facebookresearch/purplellama/blob/HEAD/CybersecurityBenchmarks/benchmark/arvo_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"744a6e2320d37acf"}},{"code_sha256_prefix":"15ff5c1284e75ee1","entry":"get_least_common_ancestor","repo":"facebookresearch/purplellama","repo_kind":"official","path":"CybersecurityBenchmarks/benchmark/arvo_utils.py","file_url":"https://github.com/facebookresearch/purplellama/blob/HEAD/CybersecurityBenchmarks/benchmark/arvo_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"15ff5c1284e75ee1"}},{"code_sha256_prefix":"af3883851ddae55c","entry":"get_file_extension","repo":"facebookresearch/purplellama","repo_kind":"official","path":"CodeShield/insecure_code_detector/languages.py","file_url":"https://github.com/facebookresearch/purplellama/blob/HEAD/CodeShield/insecure_code_detector/languages.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"af3883851ddae55c"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}