Papers › Is it worth it? Comparing six deep and classical methods for unsupervised anomaly...

Is it worth it? Comparing six deep and classical methods for unsupervised anomaly detection in time series

21 Dec 2022arXiv:2212.11080archive 2025-07-28

Ferdinand Rewicki, Joachim Denzler, Julia Niebling

Detecting anomalies in time series data is important in a variety of fields, including system monitoring, healthcare, and cybersecurity. While the abundance of available methods makes it difficult to choose the most appropriate method for a given application, each method has its strengths in detecting certain types of anomalies. In this study, we compare six unsupervised anomaly detection methods of varying complexity to determine whether more complex methods generally perform better and if certain methods are better suited to certain types of anomalies. We evaluated the methods using the UCR anomaly archive, a recent benchmark dataset for anomaly detection. We analyzed the results on a dataset and anomaly type level after adjusting the necessary hyperparameters for each method. Additionally, we assessed the ability of each method to incorporate prior knowledge about anomalies and examined the differences between point-wise and sequence-wise features. Our experiments show that classical machine learning methods generally outperform deep learning methods across a range of anomaly types.

PaperPDFCode

Code

gitlab.com/dlr-dw/is-it-worth-it-benchmark officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Anomaly DetectionTime SeriesTime Series AnalysisUnsupervised Anomaly Detection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Anomaly Detection UCR Anomaly Archive MERLIN AUC ROC 0.51 ± 0.0 #5 of 24 Archive leaderboard report
Anomaly Detection UCR Anomaly Archive MERLIN Average F1 0.27 ±0.0 #5 of 24 Archive leaderboard report
Anomaly Detection UCR Anomaly Archive Maximally Divergent Intervals (MDI) AUC ROC 0.66 ± 0.0 #6 of 24 Archive leaderboard report
Anomaly Detection UCR Anomaly Archive Maximally Divergent Intervals (MDI) Average F1 0.25 ±0.0 #6 of 24 Archive leaderboard report
Anomaly Detection UCR Anomaly Archive Graph Augmented Normalizing Flows (GANF) AUC ROC 0.63 ±0.009 #7 of 24 Archive leaderboard report
Anomaly Detection UCR Anomaly Archive Graph Augmented Normalizing Flows (GANF) Average F1 0.23 ±0.021 #7 of 24 Archive leaderboard report
Anomaly Detection UCR Anomaly Archive Transformer Network for Anomaly Detection (TranAD) AUC ROC 0.56 ±0.003 #8 of 24 Archive leaderboard report
Anomaly Detection UCR Anomaly Archive Transformer Network for Anomaly Detection (TranAD) Average F1 0.18 ±0.003 #8 of 24 Archive leaderboard report
Anomaly Detection UCR Anomaly Archive Autoencoder (AE) AUC ROC 0.58 ±0.01 #10 of 24 Archive leaderboard report
Anomaly Detection UCR Anomaly Archive Autoencoder (AE) Average F1 0.16 ± 0.013 #10 of 24 Archive leaderboard report
Anomaly Detection UCR Anomaly Archive Robust Random Cut Forest (RRCF) AUC ROC 0.56 ± 0.0019 #12 of 24 Archive leaderboard report
Anomaly Detection UCR Anomaly Archive Robust Random Cut Forest (RRCF) Average F1 0.07 ±0.011 #12 of 24 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections