{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/towards-real-time-country-level-location","title":"Towards Real-Time, Country-Level Location Classification of Worldwide Tweets","arxiv_id":"1604.07236","date":"2016-04-25","proceeding":null,"authors":["Arkaitz Zubiaga","Alex Voss","Rob Procter","Maria Liakata","Bo wang","Adam Tsakalidis"],"abstract":"In contrast to much previous work that has focused on location classification\nof tweets restricted to a specific country, here we undertake the task in a\nbroader context by classifying global tweets at the country level, which is so\nfar unexplored in a real-time scenario. We analyse the extent to which a\ntweet's country of origin can be determined by making use of eight\ntweet-inherent features for classification. Furthermore, we use two datasets,\ncollected a year apart from each other, to analyse the extent to which a model\ntrained from historical tweets can still be leveraged for classification of new\ntweets. With classification experiments on all 217 countries in our datasets,\nas well as on the top 25 countries, we offer some insights into the best use of\ntweet-inherent features for an accurate country-level classification of tweets.\nWe find that the use of a single feature, such as the use of tweet content\nalone -- the most widely used feature in previous work -- leaves much to be\ndesired. Choosing an appropriate combination of both tweet content and metadata\ncan actually lead to substantial improvements of between 20\\% and 50\\%. We\nobserve that tweet content, the user's self-reported location and the user's\nreal name, all of which are inherent in a tweet and available in a real-time\nscenario, are particularly useful to determine the country of origin. We also\nexperiment on the applicability of a model trained on historical tweets to\nclassify new tweets, finding that the choice of a particular combination of\nfeatures whose utility does not fade over time can actually lead to comparable\nperformance, avoiding the need to retrain. However, the difficulty of achieving\naccurate classification increases slightly for countries with multiple\ncommonalities, especially for English and Spanish speaking countries.","url_abs":"http://arxiv.org/abs/1604.07236v3","url_pdf":"http://arxiv.org/pdf/1604.07236v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"towards-real-time-country-level-location","repo_url":"https://github.com/MALHARULHAS/A-Country_level-location-classification-system-for-twitter-tweets-from-the-whole-world","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}