{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/extracting-and-analyzing-semantic-relatedness","title":"Extracting and Analyzing Semantic Relatedness between Cities Using News Articles","arxiv_id":"1809.02823","date":"2018-09-08","proceeding":null,"authors":["Yingjie Hu","Xinyue Ye","Shih-Lung Shaw"],"abstract":"News articles capture a variety of topics about our society. They reflect not\nonly the socioeconomic activities that happened in our physical world, but also\nsome of the cultures, human interests, and public concerns that exist only in\nthe perceptions of people. Cities are frequently mentioned in news articles,\nand two or more cities may co-occur in the same article. Such co-occurrence\noften suggests certain relatedness between the mentioned cities, and the\nrelatedness may be under different topics depending on the contents of the news\narticles. We consider the relatedness under different topics as semantic\nrelatedness. By reading news articles, one can grasp the general semantic\nrelatedness between cities, yet, given hundreds of thousands of news articles,\nit is very difficult, if not impossible, for anyone to manually read them. This\npaper proposes a computational framework which can \"read\" a large number of\nnews articles and extract the semantic relatedness between cities. This\nframework is based on a natural language processing model and employs a machine\nlearning process to identify the main topics of news articles. We describe the\noverall structure of this framework and its individual modules, and then apply\nit to an experimental dataset with more than 500,000 news articles covering the\ntop 100 U.S. cities spanning a 10-year period. We perform exploratory\nvisualization of the extracted semantic relatedness under different topics and\nover multiple years. We also analyze the impact of geographic distance on\nsemantic relatedness and find varied distance decay effects. The proposed\nframework can be used to support large-scale content analysis in city network\nresearch.","url_abs":"http://arxiv.org/abs/1809.02823v1","url_pdf":"http://arxiv.org/pdf/1809.02823v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"extracting-and-analyzing-semantic-relatedness","repo_url":"https://github.com/YingjieHu/CityRelatednessViaNews","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"articles","task_name":"Articles"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1809.02823","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}