{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/finding-structure-in-text-genome-and-other","title":"Finding Structure in Text, Genome and Other Symbolic Sequences","arxiv_id":"1207.1847","date":"2012-07-08","proceeding":null,"authors":["Ted Dunning"],"abstract":"The statistical methods derived and described in this thesis provide new ways\nto elucidate the structural properties of text and other symbolic sequences.\nGenerically, these methods allow detection of a difference in the frequency of\na single feature, the detection of a difference between the frequencies of an\nensemble of features and the attribution of the source of a text. These three\nabstract tasks suffice to solve problems in a wide variety of settings.\nFurthermore, the techniques described in this thesis can be extended to provide\na wide range of additional tests beyond the ones described here.\n  A variety of applications for these methods are examined in detail. These\napplications are drawn from the area of text analysis and genetic sequence\nanalysis. The textually oriented tasks include finding interesting collocations\nand cooccurent phrases, language identification, and information retrieval. The\nbiologically oriented tasks include species identification and the discovery of\npreviously unreported long range structure in genes. In the applications\nreported here where direct comparison is possible, the performance of these new\nmethods substantially exceeds the state of the art.\n  Overall, the methods described here provide new and effective ways to analyse\ntext and other symbolic sequences. Their particular strength is that they deal\nwell with situations where relatively little data are available. Since these\nmethods are abstract in nature, they can be applied in novel situations with\nrelative ease.","url_abs":"http://arxiv.org/abs/1207.1847v1","url_pdf":"http://arxiv.org/pdf/1207.1847v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"finding-structure-in-text-genome-and-other","repo_url":"https://github.com/rn123/japanese_text_analysis","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"information-retrieval","task_name":"Information Retrieval"},{"task_slug":"language-identification","task_name":"Language Identification"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}