{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/nationality-classification-using-name","title":"Nationality Classification Using Name Embeddings","arxiv_id":"1708.07903","date":"2017-08-25","proceeding":null,"authors":["Junting Ye","Shuchu Han","Yifan Hu","Baris Coskun","Meizhu Liu","Hong Qin","Steven Skiena"],"abstract":"Nationality identification unlocks important demographic information, with\nmany applications in biomedical and sociological research. Existing name-based\nnationality classifiers use name substrings as features and are trained on\nsmall, unrepresentative sets of labeled names, typically extracted from\nWikipedia. As a result, these methods achieve limited performance and cannot\nsupport fine-grained classification.\n  We exploit the phenomena of homophily in communication patterns to learn name\nembeddings, a new representation that encodes gender, ethnicity, and\nnationality which is readily applicable to building classifiers and other\nsystems. Through our analysis of 57M contact lists from a major Internet\ncompany, we are able to design a fine-grained nationality classifier covering\n39 groups representing over 90% of the world population. In an evaluation\nagainst other published systems over 13 common classes, our F1 score (0.795) is\nsubstantial better than our closest competitor Ethnea (0.580). To the best of\nour knowledge, this is the most accurate, fine-grained nationality classifier\navailable.\n  As a social media application, we apply our classifiers to the followers of\nmajor Twitter celebrities over six different domains. We demonstrate stark\ndifferences in the ethnicities of the followers of Trump and Obama, and in the\nsports and entertainments favored by different groups. Finally, we identify an\nanomalous political figure whose presumably inflated following appears largely\nincapable of reading the language he posts in.","url_abs":"http://arxiv.org/abs/1708.07903v1","url_pdf":"http://arxiv.org/pdf/1708.07903v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"nationality-classification-using-name","repo_url":"https://github.com/AlokElashoff/Ethnicity_Classification_URAP","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}