{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/evaluating-overfit-and-underfit-in-models-of","title":"Evaluating Overfit and Underfit in Models of Network Community Structure","arxiv_id":"1802.10582","date":"2018-02-28","proceeding":null,"authors":["Amir Ghasemian","Homa Hosseinmardi","Aaron Clauset"],"abstract":"A common data mining task on networks is community detection, which seeks an\nunsupervised decomposition of a network into structural groups based on\nstatistical regularities in the network's connectivity. Although many methods\nexist, the No Free Lunch theorem for community detection implies that each\nmakes some kind of tradeoff, and no algorithm can be optimal on all inputs.\nThus, different algorithms will over or underfit on different inputs, finding\nmore, fewer, or just different communities than is optimal, and evaluation\nmethods that use a metadata partition as a ground truth will produce misleading\nconclusions about general accuracy. Here, we present a broad evaluation of over\nand underfitting in community detection, comparing the behavior of 16\nstate-of-the-art community detection algorithms on a novel and structurally\ndiverse corpus of 406 real-world networks. We find that (i) algorithms vary\nwidely both in the number of communities they find and in their corresponding\ncomposition, given the same input, (ii) algorithms can be clustered into\ndistinct high-level groups based on similarities of their outputs on real-world\nnetworks, and (iii) these differences induce wide variation in accuracy on link\nprediction and link description tasks. We introduce a new diagnostic for\nevaluating overfitting and underfitting in practice, and use it to roughly\ndivide community detection methods into general and specialized learning\nalgorithms. Across methods and inputs, Bayesian techniques based on the\nstochastic block model and a minimum description length approach to\nregularization represent the best general learning approach, but can be\noutperformed under specific circumstances. These results introduce both a\ntheoretically principled approach to evaluate over and underfitting in models\nof network community structure and a realistic benchmark by which new methods\nmay be evaluated and compared.","url_abs":"http://arxiv.org/abs/1802.10582v3","url_pdf":"http://arxiv.org/pdf/1802.10582v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"evaluating-overfit-and-underfit-in-models-of","repo_url":"https://github.com/AGhasemian/CommunityFitNet","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"community-detection","task_name":"Community Detection"},{"task_slug":"diagnostic","task_name":"Diagnostic"},{"task_slug":"link-prediction","task_name":"Link Prediction"},{"task_slug":"stochastic-block-model","task_name":"Stochastic Block Model"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}