{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/exploring-text-datasets-by-visualizing","title":"Exploring text datasets by visualizing relevant words","arxiv_id":"1707.05261","date":"2017-07-17","proceeding":null,"authors":["Franziska Horn","Leila Arras","Grégoire Montavon","Klaus-Robert Müller","Wojciech Samek"],"abstract":"When working with a new dataset, it is important to first explore and\nfamiliarize oneself with it, before applying any advanced machine learning\nalgorithms. However, to the best of our knowledge, no tools exist that quickly\nand reliably give insight into the contents of a selection of documents with\nrespect to what distinguishes them from other documents belonging to different\ncategories. In this paper we propose to extract `relevant words' from a\ncollection of texts, which summarize the contents of documents belonging to a\ncertain class (or discovered cluster in the case of unlabeled datasets), and\nvisualize them in word clouds to allow for a survey of salient features at a\nglance. We compare three methods for extracting relevant words and demonstrate\nthe usefulness of the resulting word clouds by providing an overview of the\nclasses contained in a dataset of scientific publications as well as by\ndiscovering trending topics from recent New York Times article snippets.","url_abs":"http://arxiv.org/abs/1707.05261v1","url_pdf":"http://arxiv.org/pdf/1707.05261v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"exploring-text-datasets-by-visualizing","repo_url":"https://github.com/cod3licious/textcatvis","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"exploring-text-datasets-by-visualizing","repo_url":"https://github.com/acdreyer/thesis","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}