{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/exchangeable-modelling-of-relational-data","title":"Exchangeable modelling of relational data: checking sparsity, train-test splitting, and sparse exchangeable Poisson matrix factorization","arxiv_id":"1712.02311","date":"2017-12-06","proceeding":null,"authors":["Victor Veitch","Ekansh Sharma","Zacharie Naulet","Daniel M. Roy"],"abstract":"A variety of machine learning tasks---e.g., matrix factorization, topic\nmodelling, and feature allocation---can be viewed as learning the parameters of\na probability distribution over bipartite graphs. Recently, a new class of\nmodels for networks, the sparse exchangeable graphs, have been introduced to\nresolve some important pathologies of traditional approaches to statistical\nnetwork modelling; most notably, the inability to model sparsity (in the\nasymptotic sense). The present paper explains some practical insights arising\nfrom this work. We first show how to check if sparsity is relevant for\nmodelling a given (fixed size) dataset by using network subsampling to identify\na simple signature of sparsity. We discuss the implications of the (sparse)\nexchangeable subsampling theory for test-train dataset splitting; we argue\ncommon approaches can lead to biased results, and we propose a principled\nalternative. Finally, we study sparse exchangeable Poisson matrix factorization\nas a worked example. In particular, we show how to adapt mean field variational\ninference to the sparse exchangeable setting, allowing us to scale inference to\nhuge datasets.","url_abs":"http://arxiv.org/abs/1712.02311v1","url_pdf":"http://arxiv.org/pdf/1712.02311v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"exchangeable-modelling-of-relational-data","repo_url":"https://github.com/ekanshs/graphex-nnmf","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"variational-inference","task_name":"Variational Inference"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}