{"url":"/dataset/autorobust","name":"AutoRobust","full_name":null,"description_markdown":"This dataset is comprised of the dynamic analysis reports generated by CAPEv2, from both malware and goodware. We source the goodware as they do in Dambra et al. (https://arxiv.org/abs/2307.14657), where trough the community-maintained packages of Chocolatey they create a dataset that spans 2012 to 2020.\r\nThe malware are sourced from VirusTotal, namely samples of Portable Executable from 2017 - 2020 that they release for academic purposes.\r\nIn total, the dataset we assembled contains 26,200 PE samples: 8,600 (33\\%) goodware and 17,675 (67\\%) malware.\r\n\r\nAll samples were executed in a series of Windows 7 VMware virtual machines and without network connectivity due to policy/ethical constraints. These were orchestrated through CAPEv2, a maintained open-source successor of Cuckoo. This framework is widely used for analyzing and detecting potentially malicious binaries by executing them in an isolated environment and observing their behavior without risking the security of the host system. Every sample was executed for 150 seconds, an empirical lower bound on the time required to gather the full behavior of samples.\r\n\r\nWe tried to mitigate evasive checks by running the VMwareCloak script to remove well-known artifacts introduced by VMware. Moreover, we populated the filesystem with documents and other common types of files, to resemble a legitimate desktop workstation that malware may identify as a valuable target.\r\nThe dynamic analysis output is a detailed report with information on the syscalls invoked by the binary and all the relative flags and arguments, as well as all interactions with the file system, registry, network, and other key elements of the operating system.\r\n\r\nThe final dataset contains only the dynamic analysis reports, without the original binaries.","description_withheld":null,"homepage":"https://www.kaggle.com/datasets/greimas/malware-and-goodware-dynamic-analysis-reports","introduced_date":"2024-07-18","introduced_date_note":null,"introduced_by":null,"license":{"name":"CC-BY","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Classification","url":"/task/classification-1","datasets_with_task":"/datasets/task/classification-1"},{"name":"Binary Classification","url":"/task/binary-classification","datasets_with_task":"/datasets/task/binary-classification"},{"name":"Malware Classification","url":"/task/malware-classification","datasets_with_task":"/datasets/task/malware-classification"},{"name":"Multi-class Classification","url":"/task/multi-class-classification","datasets_with_task":"/datasets/task/multi-class-classification"},{"name":"Malware Detection","url":"/task/malware-detection","datasets_with_task":"/datasets/task/malware-detection"},{"name":"Malware Family Detection","url":"/task/malware-family-detection","datasets_with_task":"/datasets/task/malware-family-detection"},{"name":"Behavioral Malware Classification","url":"/task/behavioral-malware-classification","datasets_with_task":"/datasets/task/behavioral-malware-classification"},{"name":"Malware Analysis","url":"/task/malware-analysis","datasets_with_task":"/datasets/task/malware-analysis"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["AutoRobust"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}