{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/fame-for-sale-efficient-detection-of-fake","title":"Fame for sale: efficient detection of fake Twitter followers","arxiv_id":"1509.04098","date":"2015-09-14","proceeding":null,"authors":["Stefano Cresci","Roberto Di Pietro","Marinella Petrocchi","Angelo Spognardi","Maurizio Tesconi"],"abstract":"$\\textit{Fake followers}$ are those Twitter accounts specifically created to\ninflate the number of followers of a target account. Fake followers are\ndangerous for the social platform and beyond, since they may alter concepts\nlike popularity and influence in the Twittersphere - hence impacting on\neconomy, politics, and society. In this paper, we contribute along different\ndimensions. First, we review some of the most relevant existing features and\nrules (proposed by Academia and Media) for anomalous Twitter accounts\ndetection. Second, we create a baseline dataset of verified human and fake\nfollower accounts. Such baseline dataset is publicly available to the\nscientific community. Then, we exploit the baseline dataset to train a set of\nmachine-learning classifiers built over the reviewed rules and features. Our\nresults show that most of the rules proposed by Media provide unsatisfactory\nperformance in revealing fake followers, while features proposed in the past by\nAcademia for spam detection provide good results. Building on the most\npromising features, we revise the classifiers both in terms of reduction of\noverfitting and cost for gathering the data needed to compute the features. The\nfinal result is a novel $\\textit{Class A}$ classifier, general enough to thwart\noverfitting, lightweight thanks to the usage of the less costly features, and\nstill able to correctly classify more than 95% of the accounts of the original\ntraining set. We ultimately perform an information fusion-based sensitivity\nanalysis, to assess the global sensitivity of each of the features employed by\nthe classifier. The findings reported in this paper, other than being supported\nby a thorough experimental methodology and interesting on their own, also pave\nthe way for further investigation on the novel issue of fake Twitter followers.","url_abs":"http://arxiv.org/abs/1509.04098v2","url_pdf":"http://arxiv.org/pdf/1509.04098v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"spam-detection","task_name":"Spam detection"}],"methods":[],"datasets_introduced":[{"slug":"mib-datasets","name":"MIB Dataset","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1509.04098","atlas_url":"https://app.syntology.ai/?focus=1509.04098","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}