Browse State-of-the-Art › Twitter Bot Detection
Twitter Bot Detection
10 papers with code · 2 benchmarks · 3 datasets archive 2025-07-28
Academic studies estimate that up to 15% of Twitter users are automated bot accounts [1]. The prevalence of Twitter bots coupled with the ability of some bots to give seemingly human responses has enabled these non-human accounts to garner widespread influence. Hence, detecting non-human Twitter users or automated bot accounts using machine learning techniques has become an area of interest to researchers in the last few years.
[1] https://aaai.org/ocs/index.php/ICWSM/ICWSM17/paper/view/15587
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
2 leaderboard tables shown for this task, 2 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| MGTAB (4 rows) | RGT | MGTAB: A Multi-Relational Graph-Based Twitter Account Detection Benchmark | code | — | Compare |
| MIB Dataset (1 row) | DNA String Compression - Compression Ratio | Detecting Bot Behaviour in Social Media using Digital DNA Compression | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
10 shown of 10 papers with code (16 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
25 Oct 2018 2 repositories listedWe introduce a graphical framework that (1) generalizes existing attacks in discrete domains, (2) can accommodate complex cost functions beyond p-norms, including financial cost incurred when attacking a classifier, and…
-
30 Jun 2023 1 repository listedFor datasets without graph structure, we simply replace the GNN with an MLP, which has also shown strong performance.
-
17 Jan 2023 1 repository listedThese tools employ machine learning and often achieve near perfect performance for classification on existing datasets, suggesting bot detection is accurate, reliable and fit for use in downstream applications.
-
3 Jan 2023 1 repository listedHowever, in addition to low annotation quality, existing benchmarks generally have incomplete user relationships, suppressing graph-based account detection research.
-
17 Aug 2022 1 repository listedIn addition, given the stealing behavior of novel Twitter bots, BIC proposes to model semantic consistency in tweets based on attention weights while using it to augment the decision process.
-
9 Jun 2022 1 repository listed Syntology ran 1 of 4 samples · 3 unverifiedTwitter bot detection has become an increasingly important task to combat misinformation, facilitate social media moderation, and preserve the integrity of the online discourse.
-
8 Dec 2021 1 repository listedTwitter is one of the most popular social networks attracting millions of users, while a considerable proportion of online discourse is captured.
-
11 May 2020 1 repository listedThis paper presents state of the art methods for addressing three important challenges in automated fake news detection: fake news detection, domain identification, and bot identification in tweets.
-
5 Dec 2019 1 repository listedIn our approach, we employ a lossless compression algorithm on these Digital DNA sequences and use the compression statistics as a measure of predictability in the behaviour of a group of Twitter accounts.
-
1 Jul 2019 1 repository listedIn this work, we present our approach for the Author Profiling task of PAN 2019.
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections