{"url":"/dataset/fft-75","name":"FFT-75","full_name":null,"description_markdown":"The **FFT-75** dataset contains randomly sampled, potentially overlapping file fragments from 75 popular file types. It is a diverse and balanced dataset which is labeled with class IDs and is ready for training supervised machine learning models. We distinguish 6 different scenarios with different granularity and provide variants with 512 and 4096-byte blocks. In each case, we sampled a balanced dataset and split the data as follows: 80% for training, 10% for testing and 10% for validation.","description_withheld":null,"homepage":"https://ieee-dataport.org/open-access/file-fragment-type-fft-75-dataset","introduced_date":"2019-08-16","introduced_date_note":null,"introduced_by":{"paper":"/paper/fifty-large-scale-file-fragment-type","title":"FiFTy: Large-scale File Fragment Type Identification using Neural Networks","first_author":null,"url":null},"license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"modalities":[],"tasks":[],"languages":[],"variants":["FFT-75"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}