{"url":"/dataset/isadetect-dataset","name":"ISAdetect dataset","full_name":"ISAdetect binary file and object code dataset","description_markdown":"This repository holds two datasets: one with both the original binaries and the code sections extracted from them (“full dataset”), and one with only the code sections (“only code sections”). The code sections were extracted by carving out sections of the binary that were marked as executable. The binaries were scraped from Debian repositories.\r\n\r\nThere are also two CSV files available, one with full binaries and one with only code sections, which include the 293 features extracted from about 3000 binaries per architecture. These features can be used to train classifiers.\r\n\r\nThe dataset consists of thousands of binaries for the following 23 architectures: alpha, amd64, arm64, armel, armhf, hppa, i386, ia64, m68k, mips, mips64el, mipsel, powerpc, powerpcspe, powerpc64, powerpc64el, riscv, s390, s390x, sh4, sparc, sparc64 and x32.\r\n\r\nThere are 98 500 binary files, about 27 gigabytes (uncompressed) of binary files and about 15 gigabytes (uncompressed) of only code sections from those binary files.\r\n\r\nBoth datasets hold the binaries in directories named by the architecture. The files inside the folders are named as MD5 hashes of the original binary files, and a hash file ending with “.code” contains only the concatenation of all code sections of the original binary file. Each architecture folder also holds a JSON file named after the architecture, e.g. amd64 holds amd64.json. The structure of the JSON file is as follows (described in a JSON Schema-like notation)\r\n\r\nThis work is based on work by John Clemens, 2015, “Automatic classification of object code using machine learning” and De Nicolao, Pietro et al., 2018, “ELISA: ELiciting ISA of Raw Binaries for Fine-Grained Code and Data Separation”\r\n\r\nThis dataset is released as part of the following papers:\r\n\r\nSami Kairajärvi, Andrei Costin, and Timo Hämäläinen. 2020. ISAdetect: Usable automated detection of ISA (CPU architecture and endianness) for executable binary files and object code. In Tenth ACM Conference on Data and Application Security and Privacy (CODASPY’20), March 16–18, 2020, New Orleans, LA, USA. ACM, New York, NY, USA, 5 pages. https://doi.org/10.1145/3374664.3375742\r\n\r\nKairajärvi, Sami, Andrei Costin, and Timo Hämäläinen. \"Towards usable automated detection of CPU architecture and endianness for arbitrary binary files and object code sequences.\" arXiv preprint arXiv:1908.05459 (2019).\r\n\r\nKairajärvi, Sami. \"Automatic identification of architecture and endianness using binary file contents.\" (2019).\r\n\r\nThe code associated with this dataset can be found at https://github.com/kairis/isadetect\r\n\r\nChangelog: version 6 - 29.3.2020\r\n\r\n    Add Weka models\r\n\r\nversion 5 - 17.1.2020\r\n\r\n    Clean up dataset\r\n\r\nversion 4 - 13.1.2020\r\n\r\n    Initial release","description_withheld":null,"homepage":"https://etsin.fairdata.fi/dataset/80fa69af-addb-4f9a-b45c-c16011bae366","introduced_date":"2020-01-20","introduced_date_note":null,"introduced_by":{"paper":"/paper/towards-usable-automated-detection-of-cpu","title":"Towards usable automated detection of CPU architecture and endianness for arbitrary binary files and object code sequences","first_author":null,"url":null},"license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"modalities":[],"tasks":[{"name":"Code Search","url":"/task/code-search","datasets_with_task":"/datasets/task/code-search"},{"name":"Annotated Code Search","url":"/task/annotated-code-search","datasets_with_task":"/datasets/task/annotated-code-search"},{"name":"Code Classification","url":"/task/code-classification","datasets_with_task":"/datasets/task/code-classification"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["ISAdetect dataset"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}