{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/outlier-detection-on-mixed-type-data-an","title":"Outlier Detection on Mixed-Type Data: An Energy-based Approach","arxiv_id":"1608.04830","date":"2016-08-17","proceeding":null,"authors":["Kien Do","Truyen Tran","Dinh Phung","Svetha Venkatesh"],"abstract":"Outlier detection amounts to finding data points that differ significantly\nfrom the norm. Classic outlier detection methods are largely designed for\nsingle data type such as continuous or discrete. However, real world data is\nincreasingly heterogeneous, where a data point can have both discrete and\ncontinuous attributes. Handling mixed-type data in a disciplined way remains a\ngreat challenge. In this paper, we propose a new unsupervised outlier detection\nmethod for mixed-type data based on Mixed-variate Restricted Boltzmann Machine\n(Mv.RBM). The Mv.RBM is a principled probabilistic method that models data\ndensity. We propose to use \\emph{free-energy} derived from Mv.RBM as outlier\nscore to detect outliers as those data points lying in low density regions. The\nmethod is fast to learn and compute, is scalable to massive datasets. At the\nsame time, the outlier score is identical to data negative log-density up-to an\nadditive constant. We evaluate the proposed method on synthetic and real-world\ndatasets and demonstrate that (a) a proper handling mixed-types is necessary in\noutlier detection, and (b) free-energy of Mv.RBM is a powerful and efficient\noutlier scoring method, which is highly competitive against state-of-the-arts.","url_abs":"http://arxiv.org/abs/1608.04830v1","url_pdf":"http://arxiv.org/pdf/1608.04830v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"outlier-detection-on-mixed-type-data-an","repo_url":"https://github.com/clarken92/Mixed-variate-RBM","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"outlier-detection","task_name":"Outlier Detection"},{"task_slug":"type","task_name":"Vocal Bursts Type Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}