{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/finding-statistically-significant-attribute","title":"Finding Statistically Significant Attribute Interactions","arxiv_id":"1612.07597","date":"2016-12-22","proceeding":null,"authors":["Andreas Henelius","Antti Ukkonen","Kai Puolamäki"],"abstract":"In many data exploration tasks it is meaningful to identify groups of\nattribute interactions that are specific to a variable of interest. For\ninstance, in a dataset where the attributes are medical markers and the\nvariable of interest (class variable) is binary indicating presence/absence of\ndisease, we would like to know which medical markers interact with respect to\nthe binary class label. These interactions are useful in several practical\napplications, for example, to gain insight into the structure of the data, in\nfeature selection, and in data anonymisation. We present a novel method, based\non statistical significance testing, that can be used to verify if the data set\nhas been created by a given factorised class-conditional joint distribution,\nwhere the distribution is parametrised by a partition of its attributes.\nFurthermore, we provide a method, named ASTRID, for automatically finding a\npartition of attributes describing the distribution that has generated the\ndata. State-of-the-art classifiers are utilised to capture the interactions\npresent in the data by systematically breaking attribute interactions and\nobserving the effect of this breaking on classifier performance. We empirically\ndemonstrate the utility of the proposed method with examples using real and\nsynthetic data.","url_abs":"http://arxiv.org/abs/1612.07597v2","url_pdf":"http://arxiv.org/pdf/1612.07597v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"finding-statistically-significant-attribute","repo_url":"https://github.com/bwrc/astrid-r","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"feature-selection","task_name":"feature selection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}