{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/building-networks-of-shared-research","title":"Building networks of shared research interests by embedding words into a representation space","arxiv_id":"2502.07042","date":"2025-02-10","proceeding":null,"authors":["Art Poon"],"abstract":"Departments within a university are not only administrative units, but also an effort to gather investigators around common fields of academic study. A pervasive challenge is connecting members with shared research interests both within and between departments. Here I describe a workflow that adapts methods from natural language processing to generate a network connecting $n=79$ members of a university department, or multiple departments within a faculty ($n=278$), based on common topics in their research publications. After extracting and processing terms from $n=16,901$ abstracts in the PubMed database, the co-occurrence of terms is encoded in a sparse document-term matrix. Based on the angular distances between the presence-absence vectors for every pair of terms, I use the uniform manifold approximation and projection (UMAP) method to embed the terms into a representational space such that terms that tend to appear in the same documents are closer together. Each author's corpus defines a probability distribution over terms in this space. Using the Wasserstein distance to quantify the similarity between these distributions, I generate a distance matrix among authors that can be analyzed and visualized as a graph. I demonstrate that this nonparametric method produces clusters with distinct themes that are consistent with some academic divisions, while identifying untapped connections among members. A documented workflow comprising Python and R scripts is available under the MIT license at https://github.com/PoonLab/tragula.","url_abs":"https://arxiv.org/abs/2502.07042v1","url_pdf":"https://arxiv.org/pdf/2502.07042v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"links_only","authors_date_abstract":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license), from the Kaggle arXiv metadata snapshot of 2026-09-12"},"code_links":[{"paper_slug":"building-networks-of-shared-research","repo_url":"https://github.com/poonlab/tragula","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}