{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/how-many-topics-stability-analysis-for-topic","title":"How Many Topics? Stability Analysis for Topic Models","arxiv_id":"1404.4606","date":"2014-04-16","proceeding":null,"authors":["Derek Greene","Derek O'Callaghan","Pádraig Cunningham"],"abstract":"Topic modeling refers to the task of discovering the underlying thematic\nstructure in a text corpus, where the output is commonly presented as a report\nof the top terms appearing in each topic. Despite the diversity of topic\nmodeling algorithms that have been proposed, a common challenge in successfully\napplying these techniques is the selection of an appropriate number of topics\nfor a given corpus. Choosing too few topics will produce results that are\noverly broad, while choosing too many will result in the \"over-clustering\" of a\ncorpus into many small, highly-similar topics. In this paper, we propose a\nterm-centric stability analysis strategy to address this issue, the idea being\nthat a model with an appropriate number of topics will be more robust to\nperturbations in the data. Using a topic modeling approach based on matrix\nfactorization, evaluations performed on a range of corpora show that this\nstrategy can successfully guide the model selection process.","url_abs":"http://arxiv.org/abs/1404.4606v3","url_pdf":"http://arxiv.org/pdf/1404.4606v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"how-many-topics-stability-analysis-for-topic","repo_url":"https://github.com/derekgreene/topic-stability","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"how-many-topics-stability-analysis-for-topic","repo_url":"https://github.com/machine-intelligence-laboratory/OptimalNumberOfTopics","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"model-selection","task_name":"Model Selection"},{"task_slug":"topic-models","task_name":"Topic Models"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1404.4606","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}