{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-review-of-features-for-the-discrimination","title":"A Review of Features for the Discrimination of Twitter Users: Application to the Prediction of Offline Influence","arxiv_id":"1509.06585","date":"2015-09-22","proceeding":null,"authors":["Jean-Valère Cossu","Vincent Labatut","Nicolas Dugué"],"abstract":"Many works related to Twitter aim at characterizing its users in some way:\nrole on the service (spammers, bots, organizations, etc.), nature of the user\n(socio-professional category, age, etc.), topics of interest , and others.\nHowever, for a given user classification problem, it is very difficult to\nselect a set of appropriate features, because the many features described in\nthe literature are very heterogeneous, with name overlaps and collisions, and\nnumerous very close variants. In this article, we review a wide range of such\nfeatures. In order to present a clear state-of-the-art description, we unify\ntheir names, definitions and relationships, and we propose a new, neutral,\ntypology. We then illustrate the interest of our review by applying a selection\nof these features to the offline influence detection problem. This task\nconsists in identifying users which are influential in real-life, based on\ntheir Twitter account and related data. We show that most features deemed\nefficient to predict online influence, such as the numbers of retweets and\nfollowers, are not relevant to this problem. However, We propose several\ncontent-based approaches to label Twitter users as Influencers or not. We also\nrank them according to a predicted influence level. Our proposals are evaluated\nover the CLEF RepLab 2014 dataset, and outmatch state-of-the-art methods.","url_abs":"http://arxiv.org/abs/1509.06585v3","url_pdf":"http://arxiv.org/pdf/1509.06585v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-review-of-features-for-the-discrimination","repo_url":"https://github.com/CompNet/Influence","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}