{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-u-statistic-approach-to-hypothesis-testing","title":"A U-statistic Approach to Hypothesis Testing for Structure Discovery in Undirected Graphical Models","arxiv_id":"1604.01733","date":"2016-04-06","proceeding":null,"authors":["Wacha Bounliphone","Matthew Blaschko"],"abstract":"Structure discovery in graphical models is the determination of the topology\nof a graph that encodes conditional independence properties of the joint\ndistribution of all variables in the model. For some class of probability\ndistributions, an edge between two variables is present if and only if the\ncorresponding entry in the precision matrix is non-zero. For a finite sample\nestimate of the precision matrix, entries close to zero may be due to low\nsample effects, or due to an actual association between variables; these two\ncases are not readily distinguishable. %Fisher provided a hypothesis test based\non a parametric approximation to the distribution of an entry in the precision\nmatrix of a Gaussian distribution, but this may not provide valid upper bounds\non $p$-values for non-Gaussian distributions. Many related works on this topic\nconsider potentially restrictive distributional or sparsity assumptions that\nmay not apply to a data sample of interest, and direct estimation of the\nuncertainty of an estimate of the precision matrix for general distributions\nremains challenging. Consequently, we make use of results for $U$-statistics\nand apply them to the covariance matrix. By probabilistically bounding the\ndistortion of the covariance matrix, we can apply Weyl's theorem to bound the\ndistortion of the precision matrix, yielding a conservative, but sound test\nthreshold for a much wider class of distributions than considered in previous\nworks. The resulting test enables one to answer with statistical significance\nwhether an edge is present in the graph, and convergence results are known for\na wide range of distributions. The computational complexities is linear in the\nsample size enabling the application of the test to large data samples for\nwhich computation time becomes a limiting factor. We experimentally validate\nthe correctness and scalability of the test on multivariate distributions for\nwhich the distributional assumptions of competing tests result in\nunderestimates of the false positive ratio. By contrast, the proposed test\nremains sound, promising to be a useful tool for hypothesis testing for diverse\nreal-world problems.","url_abs":"http://arxiv.org/abs/1604.01733v1","url_pdf":"http://arxiv.org/pdf/1604.01733v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-u-statistic-approach-to-hypothesis-testing","repo_url":"https://github.com/wbounliphone/Ustatistics_Approach_For_SD","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"hypothesis-testing","task_name":"Two-sample testing"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}