Papers › Using Convolutional Neural Networks for Classification of Malware represented as Images

Using Convolutional Neural Networks for Classification of Malware represented as Images

27 Aug 2018archive 2025-07-28

Daniel Gibert, Carles Mateu, Jordi Planes & Ramon Vicens

The number of malicious files detected every year are counted by millions. One of the main reasons for these high volumes of different files is the fact that, in order to evade detection, malware authors add mutation. This means that malicious files belonging to the same family, with the same malicious behavior, are constantly modified or obfuscated using several techniques, in such a way that they look like different files. In order to be effective in analyzing and classifying such large amounts of files, we need to be able to categorize them into groups and identify their respective families on the basis of their behavior. In this paper, malicious software is visualized as gray scale images since its ability to capture minor changes while retaining the global structure helps to detect variations. Motivated by the visual similarity between malware samples of the same family, we propose a file agnostic deep learning approach for malware categorization to efficiently group malicious software into families based on a set of discriminant patterns extracted from their visualization as images. The suitability of our approach is evaluated against two benchmarks: the MalImg dataset and the Microsoft Malware Classification Challenge dataset. Experimental comparison demonstrates its superior performance with respect to state-of-the-art techniques.

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

General ClassificationMalware Classification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Malware Classification Malimg Dataset Gray-scale IMG CNN Accuracy (10-fold) 0.9848 #1 of 5 Archive leaderboard report
Malware Classification Malimg Dataset Gray-scale IMG CNN Macro F1 (10-fold) 0.9580 #1 of 5 Archive leaderboard report
Malware Classification Microsoft Malware Classification Challenge Gray-scale IMG CNN Accuracy (10-fold) 0.9750 #16 of 29 Archive leaderboard report
Malware Classification Microsoft Malware Classification Challenge Gray-scale IMG CNN Accuracy (5-fold) 0.973 #16 of 29 Archive leaderboard report
Malware Classification Microsoft Malware Classification Challenge Gray-scale IMG CNN LogLoss 0.184483 #16 of 29 Archive leaderboard report
Malware Classification Microsoft Malware Classification Challenge Gray-scale IMG CNN Macro F1 (10-fold) 0.9400 #16 of 29 Archive leaderboard report
Malware Classification Microsoft Malware Classification Challenge Haralick features + XGBoost Accuracy (5-fold) 0.9550 #27 of 29 Archive leaderboard report
Malware Classification Microsoft Malware Classification Challenge LBP features + XGBoost Accuracy (5-fold) 0.951 #28 of 29 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections