{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/clustering-based-feature-learning-on-variable","title":"Clustering Based Feature Learning on Variable Stars","arxiv_id":"1602.08977","date":"2016-02-29","proceeding":null,"authors":["Cristóbal Mackenzie","Karim Pichara","Pavlos Protopapas"],"abstract":"The success of automatic classification of variable stars strongly depends on\nthe lightcurve representation. Usually, lightcurves are represented as a vector\nof many statistical descriptors designed by astronomers called features. These\ndescriptors commonly demand significant computational power to calculate,\nrequire substantial research effort to develop and do not guarantee good\nperformance on the final classification task. Today, lightcurve representation\nis not entirely automatic; algorithms that extract lightcurve features are\ndesigned by humans and must be manually tuned up for every survey. The vast\namounts of data that will be generated in future surveys like LSST mean\nastronomers must develop analysis pipelines that are both scalable and\nautomated. Recently, substantial efforts have been made in the machine learning\ncommunity to develop methods that prescind from expert-designed and manually\ntuned features for features that are automatically learned from data. In this\nwork we present what is, to our knowledge, the first unsupervised feature\nlearning algorithm designed for variable stars. Our method first extracts a\nlarge number of lightcurve subsequences from a given set of photometric data,\nwhich are then clustered to find common local patterns in the time series.\nRepresentatives of these patterns, called exemplars, are then used to transform\nlightcurves of a labeled set into a new representation that can then be used to\ntrain an automatic classifier. The proposed algorithm learns the features from\nboth labeled and unlabeled lightcurves, overcoming the bias generated when the\nlearning process is done only with labeled data. We test our method on MACHO\nand OGLE datasets; the results show that the classification performance we\nachieve is as good and in some cases better than the performance achieved using\ntraditional features, while the computational cost is significantly lower.","url_abs":"http://arxiv.org/abs/1602.08977v1","url_pdf":"http://arxiv.org/pdf/1602.08977v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"clustering-based-feature-learning-on-variable","repo_url":"https://github.com/cmackenziek/tsfl","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"classification-of-variable-stars","task_name":"Classification Of Variable Stars"},{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"time-series","task_name":"Time Series Analysis"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}