{"url":"/task/vulnerability-detection","name":"Vulnerability Detection","slug":"vulnerability-detection","description_markdown":"Vulnerability detection plays a crucial role in safeguarding against these threats by identifying weaknesses and potential entry points that malicious actors could exploit. Through advanced scanning techniques and penetration testing, vulnerability detection tools meticulously analyze web applications and websites for vulnerabilities such as SQL injection, cross-site scripting (XSS), and insecure authentication mechanisms.\r\n\r\nBy proactively identifying and addressing vulnerabilities, organizations can strengthen their online security posture and mitigate the risk of data breaches, financial loss, and reputational damage. Additionally, vulnerability detection empowers businesses to stay compliant with industry regulations and standards, demonstrating their commitment to safeguarding sensitive information and maintaining the trust of their customers. With the evolving threat landscape and increasingly sophisticated attack vectors, investing in robust vulnerability detection measures is paramount for staying one step ahead of cyber threats and ensuring the resilience of web-based platforms and services.","categories":[{"name":"Miscellaneous","url":"/area/miscellaneous"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":216,"papers_with_code":72,"benchmarks":2,"benchmark_tables_in_archive":2,"benchmark_tables_shown":2,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":7,"subtasks":0,"parent_tasks":0},"benchmarks":[{"leaderboard":"/sota/vulnerability-detection-on-vulscriber","slug":"vulnerability-detection-on-vulscriber","dataset":"VulScribeR","dataset_url":"/dataset/vulscriber","rows_in_archive":6,"metrics":["F1 Score"],"first_row_in_archive_order":{"model":"Reveal Model - Tested on Reveal (Training on Devign + VulScribeR 20K + Extra Cleans)","paper_title":"VulScribeR: Exploring RAG-based Vulnerability Augmentation with LLMs","paper_url":"/paper/exploring-rag-based-vulnerability","paper_date":"2024-08-07","arxiv_id":"2408.04125","code_links":[{"title":"VulScribeR/VulScribeR","url":"https://github.com/VulScribeR/VulScribeR"}],"syntology":null}},{"leaderboard":"/sota/vulnerability-detection-on-vulnerability-java","slug":"vulnerability-detection-on-vulnerability-java","dataset":"Vulnerability Java Dataset","dataset_url":"/dataset/vulnerability-java-dataset","rows_in_archive":2,"metrics":["AUC","F1"],"first_row_in_archive_order":{"model":"WizardCoder","paper_title":"Finetuning Large Language Models for Vulnerability Detection","paper_url":"/paper/finetuning-large-language-models-for","paper_date":"2024-01-30","arxiv_id":"2401.17010","code_links":[{"title":"rmusab/vul-llm-finetune","url":"https://github.com/rmusab/vul-llm-finetune"}],"syntology":null}}],"datasets":[{"url":"/dataset/cvefixes","name":"CVEfixes","full_name":"","num_papers_in_archive":38},{"url":"/dataset/castle-benchmark","name":"CASTLE Benchmark","full_name":"CASTLE Benchmark C@250","num_papers_in_archive":1},{"url":"/dataset/iotvulcode","name":"IoTvulCode","full_name":"","num_papers_in_archive":1},{"url":"/dataset/realvul","name":"RealVul","full_name":"RealVul-Vulnerability Dataset following realistic settings","num_papers_in_archive":1},{"url":"/dataset/vulnerability-java-dataset","name":"Vulnerability Java Dataset","full_name":"","num_papers_in_archive":1},{"url":"/dataset/vulnerable-verified-smart-contracts","name":"Vulnerable Verified Smart Contracts","full_name":"","num_papers_in_archive":1},{"url":"/dataset/vulscriber","name":"VulScribeR","full_name":"VulScriber: 22K+ unfiltered vul samples generated with ChatGPT via Injection","num_papers_in_archive":1}],"subtasks":[],"parent_tasks":[],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":72,"tagged_in_all":216,"items":[{"url":"/paper/nyu-ctf-dataset-a-scalable-open-source","title":"NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security","date":"2024-06-08","arxiv_id":"2406.05590","repositories_listed":5,"syntology":{"n":5,"n_ran":5,"n_unverified":0,"n_pointer_only":4}},{"url":"/paper/vuldeepecker-a-deep-learning-based-system-for","title":"VulDeePecker: A Deep Learning-Based System for Vulnerability Detection","date":"2018-01-05","arxiv_id":"1801.01681","repositories_listed":5,"syntology":{"n":1,"n_ran":0,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/sysevr-a-framework-for-using-deep-learning-to","title":"SySeVR: A Framework for Using Deep Learning to Detect Software Vulnerabilities","date":"2018-07-18","arxiv_id":"1807.06756","repositories_listed":4,"syntology":{"n":8,"n_ran":1,"n_unverified":7,"n_pointer_only":1}},{"url":"/paper/limits-of-machine-learning-for-automatic","title":"Uncovering the Limits of Machine Learning for Automatic Vulnerability Detection","date":"2023-06-28","arxiv_id":"2306.17193","repositories_listed":3,"syntology":null},{"url":"/paper/safe-self-attentive-function-embeddings-for","title":"SAFE: Self-Attentive Function Embeddings for Binary Similarity","date":"2018-11-13","arxiv_id":"1811.05296","repositories_listed":3,"syntology":{"n":4,"n_ran":0,"n_unverified":4,"n_pointer_only":1}},{"url":"/paper/deep-smart-contract-intent-detection","title":"Deep Smart Contract Intent Detection","date":"2022-11-19","arxiv_id":"2211.10724","repositories_listed":2,"syntology":null},{"url":"/paper/trex-learning-execution-semantics-from-micro","title":"Trex: Learning Execution Semantics from Micro-Traces for Binary Similarity","date":"2020-12-16","arxiv_id":"2012.08680","repositories_listed":2,"syntology":null},{"url":"/paper/automated-vulnerability-detection-in-source","title":"Automated Vulnerability Detection in Source Code Using Deep Representation Learning","date":"2018-07-11","arxiv_id":"1807.04320","repositories_listed":2,"syntology":null},{"url":"/paper/boosting-vulnerability-detection-of-llms-via","title":"Boosting Vulnerability Detection of LLMs via Curriculum Preference Optimization with Synthetic Reasoning Data","date":"2025-06-09","arxiv_id":"2506.07390","repositories_listed":1,"syntology":null},{"url":"/paper/sv-trusteval-c-evaluating-structure-and","title":"SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis","date":"2025-05-27","arxiv_id":"2505.20630","repositories_listed":1,"syntology":null},{"url":"/paper/craken-cybersecurity-llm-agent-with-knowledge","title":"CRAKEN: Cybersecurity LLM Agent with Knowledge-Based Execution","date":"2025-05-21","arxiv_id":"2505.17107","repositories_listed":1,"syntology":null},{"url":"/paper/are-sparse-autoencoders-useful-for-java","title":"Are Sparse Autoencoders Useful for Java Function Bug Detection?","date":"2025-05-15","arxiv_id":"2505.10375","repositories_listed":1,"syntology":null},{"url":"/paper/can-you-really-trust-code-copilots-evaluating","title":"Can You Really Trust Code Copilots? Evaluating Large Language Models from a Code Security Perspective","date":"2025-05-15","arxiv_id":"2505.10494","repositories_listed":1,"syntology":null},{"url":"/paper/enhancing-large-language-models-with-faster","title":"Enhancing Large Language Models with Faster Code Preprocessing for Vulnerability Detection","date":"2025-05-08","arxiv_id":"2505.05600","repositories_listed":1,"syntology":null},{"url":"/paper/program-semantic-inequivalence-game-with","title":"Program Semantic Inequivalence Game with Large Language Models","date":"2025-05-02","arxiv_id":"2505.03818","repositories_listed":1,"syntology":null},{"url":"/paper/the-hitchhiker-s-guide-to-program-analysis","title":"The Hitchhiker's Guide to Program Analysis, Part II: Deep Thoughts by LLMs","date":"2025-04-16","arxiv_id":"2504.11711","repositories_listed":1,"syntology":null},{"url":"/paper/r2vul-learning-to-reason-about-software","title":"R2Vul: Learning to Reason about Software Vulnerabilities with Reinforcement Learning and Structured Reasoning Distillation","date":"2025-04-07","arxiv_id":"2504.04699","repositories_listed":1,"syntology":null},{"url":"/paper/responsible-development-of-offensive-ai","title":"Responsible Development of Offensive AI","date":"2025-04-03","arxiv_id":"2504.02701","repositories_listed":1,"syntology":null},{"url":"/paper/reasoning-with-llms-for-zero-shot","title":"Reasoning with LLMs for Zero-Shot Vulnerability Detection","date":"2025-03-22","arxiv_id":"2503.17885","repositories_listed":1,"syntology":null},{"url":"/paper/castle-benchmarking-dataset-for-static-code","title":"CASTLE: Benchmarking Dataset for Static Code Analyzers and LLMs towards CWE Detection","date":"2025-03-12","arxiv_id":"2503.09433","repositories_listed":1,"syntology":null},{"url":"/paper/mtvhunter-smart-contracts-vulnerability","title":"MTVHunter: Smart Contracts Vulnerability Detection Based on Multi-Teacher Knowledge Translation","date":"2025-02-24","arxiv_id":"2502.16955","repositories_listed":1,"syntology":null},{"url":"/paper/how-to-select-pre-trained-code-models-for","title":"How to Select Pre-Trained Code Models for Reuse? A Learning Perspective","date":"2025-01-07","arxiv_id":"2501.03783","repositories_listed":1,"syntology":null},{"url":"/paper/investigating-large-language-models-for-code","title":"Investigating Large Language Models for Code Vulnerability Detection: An Experimental Study","date":"2024-12-24","arxiv_id":"2412.18260","repositories_listed":1,"syntology":null},{"url":"/paper/leveraging-generative-ai-to-enhance-automated","title":"Leveraging Generative AI to Enhance Automated Vulnerability Scoring","date":"2024-12-07","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/cryptoformaleval-integrating-llms-and-formal","title":"CryptoFormalEval: Integrating LLMs and Formal Verification for Automated Cryptographic Protocol Vulnerability Detection","date":"2024-11-20","arxiv_id":"2411.13627","repositories_listed":1,"syntology":null},{"url":"/paper/is-function-similarity-over-engineered","title":"Is Function Similarity Over-Engineered? Building a Benchmark","date":"2024-10-30","arxiv_id":"2410.22677","repositories_listed":1,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/vulnerability-detection-via-topological","title":"Vulnerability Detection via Topological Analysis of Attention Maps","date":"2024-10-04","arxiv_id":"2410.03470","repositories_listed":1,"syntology":null},{"url":"/paper/exploring-rag-based-vulnerability","title":"VulScribeR: Exploring RAG-based Vulnerability Augmentation with LLMs","date":"2024-08-07","arxiv_id":"2408.04125","repositories_listed":1,"syntology":null},{"url":"/paper/eatvul-chatgpt-based-evasion-attack-against","title":"EaTVul: ChatGPT-based Evasion Attack Against Software Vulnerability Detection","date":"2024-07-27","arxiv_id":"2407.19216","repositories_listed":1,"syntology":null},{"url":"/paper/eyeballvul-a-future-proof-benchmark-for","title":"eyeballvul: a future-proof benchmark for vulnerability detection in the wild","date":"2024-07-11","arxiv_id":"2407.08708","repositories_listed":1,"syntology":null}],"syntology_records":5,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}