{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sequence-graph-transform-sgt-a-feature","title":"Sequence Graph Transform (SGT): A Feature Extraction Function for Sequence Data Mining (Extended Version)","arxiv_id":"1608.03533","date":"2016-08-11","proceeding":null,"authors":["Chitta Ranjan","Samaneh Ebrahimi","Kamran Paynabar"],"abstract":"The ubiquitous presence of sequence data across fields such as the web,\nhealthcare, bioinformatics, and text mining has made sequence mining a vital\nresearch area. However, sequence mining is particularly challenging because of\ndifficulty in finding (dis)similarity/distance between sequences. This is\nbecause a distance measure between sequences is not obvious due to their\nunstructuredness---arbitrary strings of arbitrary length. Feature\nrepresentations, such as n-grams, are often used but they either compromise on\nextracting both short- and long-term sequence patterns or have a high\ncomputation. We propose a new function, Sequence Graph Transform (SGT), that\nextracts the short- and long-term sequence features and embeds them in a\nfinite-dimensional feature space. Importantly, SGT has low computation and can\nextract any amount of short- to long-term patterns without any increase in the\ncomputation, also proved theoretically in this paper. Due to this, SGT yields\nsuperior result with significantly higher accuracy and lower computation\ncompared to the existing methods. We show it via several experimentation and\nSGT's real world application for clustering, classification, search and\nvisualization as examples.","url_abs":"http://arxiv.org/abs/1608.03533v9","url_pdf":"http://arxiv.org/pdf/1608.03533v9.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sequence-graph-transform-sgt-a-feature","repo_url":"https://github.com/cran2367/sgt","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"sequence-graph-transform-sgt-a-feature","repo_url":"https://github.com/usneek/sequence-embeddings","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}