Browse State-of-the-Art › Spam detection
Spam detection
38 papers with code · 1 benchmark · 2 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Traditional and Context-specific Spam Twitter (1 row) | BERT | Traditional and context-specific spam detection in low resource settings | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
2 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
2 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 38 papers with code (117 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
2 Jun 2023 2 repositories listedGraph Anomaly Detection (GAD) is a technique used to identify abnormal nodes within graphs, finding applications in network security, fraud detection, social media spam detection, and various other domains.
-
30 Jun 2022 2 repositories listedIn Deep Convolutional Neural Networks (DCNNs), the parameter count in pointwise convolutions quickly grows due to the multiplication of the filters and input channels from the preceding layer.
-
10 Oct 2021 2 repositories listedThe increase in people's use of mobile messaging services has led to the spread of social engineering attacks like phishing, considering that spam text is one of the main factors in the dissemination of phishing attacks…
-
14 Apr 2020 2 repositories listed Syntology ran 0 of 5 samples · 5 unverifiedWe show that by applying a regularization method, which we call RIPPLe, and an initialization procedure, which we call Embedding Surgery, such attacks are possible even with limited knowledge of the dataset and…
-
3 Dec 2019 2 repositories listedThere are various costs for attackers to manipulate the features of security classifiers.
-
19 Mar 2019 2 repositories listedOnline reviews have become a vital source of information in purchasing a service (product).
-
2 Nov 2018 2 repositories listedIn this paper, we develop three attacks that can bypass a broad range of common data sanitization defenses, including anomaly detectors based on nearest neighbors, training loss, and singular-value decomposition.
-
25 Oct 2018 2 repositories listedWe introduce a graphical framework that (1) generalizes existing attacks in discrete domains, (2) can accommodate complex cost functions beyond p-norms, including financial cost incurred when attacking a classifier, and…
-
13 Jan 2018 2 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedAlthough various techniques have been proposed to generate adversarial samples for white-box attacks on text, little attention has been paid to black-box attacks, which are more realistic scenarios.
-
7 May 2025 1 repository listedTherefore, we recommend testing single classifiers and imbalance learning techniques for each new dataset and application involving imbalanced datasets as is the case in several cyber security applications.
-
5 May 2025 1 repository listedEmail spam detection is a critical task in modern communication systems, essential for maintaining productivity, security, and user experience.
-
22 Oct 2024 1 repository listedThese results emphasize the significance of URLs in search engine optimization: well-named URLs enable better topic classification, increasing the likelihood of appearing on the first page of search engine results by 4.
-
21 Jun 2024 1 repository listedSpam reviews are a pervasive problem on online platforms due to its significant impact on reputation.
-
15 Apr 2024 1 repository listedIn this study, we introduce SpamDam, a SMS spam detection framework designed to overcome key challenges in detecting and understanding SMS spam, such as the lack of public SMS spam datasets, increasing privacy concerns…
-
14 Dec 2023 1 repository listedWe aimed to find a model that can accurately detect value-expressive posts in Russian social media VKontakte.
-
22 Oct 2023 1 repository listedSecurity classifiers, designed to detect malicious content in computer systems and communications, can underperform when provided with insufficient training data.
-
19 Aug 2023 1 repository listedWe present a versatile GPU-based parallel version of Logistic Regression (LR), aiming to address the increasing demand for faster algorithms in binary classification due to large data sets.
-
24 Jul 2023 1 repository listedTo tackle the challenges posed by localized training data, we approach the problem as an out-of-distribution (OOD) data issue by by aligning the distributions between the training data, which represents a small portion…
-
29 May 2023 1 repository listed Syntology ran 3 of 4 samples · 1 unverifiedAlfred is the first system for programmatic weak supervision (PWS) that creates training data for machine learning by prompting.
-
3 Apr 2023 1 repository listedOur results demonstrate that Spam-T5 surpasses baseline models and other LLMs in the majority of scenarios, particularly when there are a limited number of training samples available.
-
21 Jul 2022 1 repository listedUsers' complex behavior can be well represented by a heterogeneous graph rich with node and edge attributes.
-
9 Jun 2022 1 repository listedThe neural network model outperforms the traditional models with an F1 score of 0.
-
4 May 2022 1 repository listedTo tackle the problem of bias towards majority classes, researchers have presented various techniques to oversample the minority class data points.
-
13 Jul 2021 1 repository listedLastly, we evaluated the detection performance of the three representation models (BoW, TFIDF and BERT) coupled with the best classifier from the baseline experiment (SVM).
-
10 May 2021 1 repository listed Syntology ran 0 of 5 samples · 5 unverifiedIn this paper, we propose to accelerate GNN inference by pruning the dimensions in each layer with negligible accuracy loss.
-
8 Apr 2021 1 repository listed Syntology ran 0 of 3 samples · 3 unverifiedData exploration is an important step of every data science and machine learning project, including those involving textual data.
-
24 Feb 2021 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedRobustness of machine learning models is critical for security related applications, where real-world adversaries are uniquely focused on evading neural network based detectors.
-
24 Dec 2020 1 repository listedOnline reviews are a vital source of information when purchasing a service or a product.
-
29 Oct 2020 1 repository listedIn this paper we perform an analytic comparison of a number of techniques used to detect fake and deceptive online reviews.
-
10 Sep 2020 1 repository listed Syntology ran 1 of 3 samples · 2 unverifiedText classification has long been a staple within Natural Language Processing (NLP) with applications spanning across diverse areas such as sentiment analysis, recommender systems and spam detection.
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections