{"url":"/dataset/draper-vdisc-dataset","name":"Draper VDisc Dataset","full_name":"reza","description_markdown":"Draper VDISC Dataset - Vulnerability Detection in Source Code\r\n\r\nThe dataset consists of the source code of 1.27 million functions mined from open source software, labeled by static analysis for potential vulnerabilities. For more details on the dataset and benchmark results, see https://arxiv.org/abs/1807.04320.\r\n\r\nThe data is provided in three HDF5 files corresponding to an 80:10:10 train/validate/test split, matching the splits used in our paper. The combined file size is roughly 1 GB. Each function's raw source code, starting from the function name, is stored as a variable-length UTF-8 string. Five binary 'vulnerability' labels are provided for each function, corresponding to the four most common CWEs in our data plus all others:\r\n\r\n* CWE-120 (3.7% of functions)\r\n* CWE-119 (1.9% of functions)\r\n* CWE-469 (0.95% of functions)\r\n* CWE-476 (0.21% of functions)\r\n* CWE-other (2.7% of functions)\r\n\r\nFunctions may have more than one detected CWE each.\r\n\r\nPlease cite our paper if you use this dataset in a publication: [https://arxiv.org/abs/1807.04320](https://arxiv.org/abs/1807.04320)","description_withheld":null,"homepage":"https://osf.io/d45bw/","introduced_date":"2018-07-11","introduced_date_note":null,"introduced_by":{"paper":"/paper/automated-vulnerability-detection-in-source","title":"Automated Vulnerability Detection in Source Code Using Deep Representation Learning","first_author":"Rebecca L. Russell","url":null},"license":{"name":"CC-By Attribution 4.0 International","url":null},"modalities":[],"tasks":[],"languages":[],"variants":["Draper VDisc Dataset"],"data_loaders":[],"num_papers_in_archive":4,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}