{"url":"/dataset/insider-threat-test-dataset","name":"Insider Threat Test Dataset","full_name":"Insider Threat Test Dataset","description_markdown":"The Insider Threat Test Dataset is a collection of synthetic insider threat test datasets that provide both background and malicious actor synthetic data.\r\n\r\n\r\n\r\nThe CERT Division, in partnership with ExactData, LLC, and under sponsorship from DARPA I2O, generated a collection of synthetic insider threat test datasets. These datasets provide both synthetic background data and data from synthetic malicious actors.\r\n\r\nFor more background on this data, please see the paper, Bridging the Gap: A Pragmatic Approach to Generating Insider Threat Data.\r\n\r\nDatasets are organized according to the data generator release that created them. Most releases include multiple datasets (e.g., r3.1 and r3.2). Generally, later releases include a superset of the data generation functionality of earlier releases. Each dataset file contains a readme file that provides detailed notes about the features of that release.\r\n\r\nThe answer key file answers.tar.bz2 contains the details of the malicious activity included in each dataset, including descriptions of the scenarios enacted and the identifiers of the synthetic users involved.","description_withheld":null,"homepage":"https://kilthub.cmu.edu/articles/dataset/Insider_Threat_Test_Dataset/12841247","introduced_date":"2020-09-30","introduced_date_note":null,"introduced_by":null,"license":{"name":"Creative Commons","url":"https://creativecommons.org/licenses/by/4.0/"},"modalities":[{"name":"Tabular","url":"/datasets/modality/tabular"}],"tasks":[{"name":"Classification","url":"/task/classification-1","datasets_with_task":"/datasets/task/classification-1"},{"name":"Anomaly Detection","url":"/task/anomaly-detection","datasets_with_task":"/datasets/task/anomaly-detection"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Insider Threat Test Dataset"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/classification-on-insider-threat-test-dataset","task":"Classification","dataset_variant":"Insider Threat Test Dataset","rows":1,"metrics":["F1 score"],"first_row_in_archive_order":{"model":"TD-CNN- Attention","paper":"/paper/attention-to-patterns-is-all-you-need-for","metrics":{"F1 score":"99.71"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/attention-to-patterns-is-all-you-need-for","title":"Attention to Patterns is all you need for Insider threat detection","date":"2024-10-26","rows_on_this_dataset":1,"code_links":0,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}