{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/data-adaptive-statistics-for-multiple","title":"Data-adaptive statistics for multiple hypothesis testing in high-dimensional settings","arxiv_id":"1704.07008","date":"2017-04-24","proceeding":null,"authors":["Weixin Cai","Nima S. Hejazi","Alan E. Hubbard"],"abstract":"Current statistical inference problems in areas like astronomy, genomics, and\nmarketing routinely involve the simultaneous testing of thousands -- even\nmillions -- of null hypotheses. For high-dimensional multivariate\ndistributions, these hypotheses may concern a wide range of parameters, with\ncomplex and unknown dependence structures among variables. In analyzing such\nhypothesis testing procedures, gains in efficiency and power can be achieved by\nperforming variable reduction on the set of hypotheses prior to testing. We\npresent in this paper an approach using data-adaptive multiple testing that\nserves exactly this purpose. This approach applies data mining techniques to\nscreen the full set of covariates on equally sized partitions of the whole\nsample via cross-validation. This generalized screening procedure is used to\ncreate average ranks for covariates, which are then used to generate a reduced\n(sub)set of hypotheses, from which we compute test statistics that are\nsubsequently subjected to standard multiple testing corrections. The principal\nadvantage of this methodology lies in its providing valid statistical inference\nwithout the \\textit{a priori} specifying which hypotheses will be tested. Here,\nwe present the theoretical details of this approach, confirm its validity via a\nsimulation study, and exemplify its use by applying it to the analysis of data\non microRNA differential expression.","url_abs":"http://arxiv.org/abs/1704.07008v1","url_pdf":"http://arxiv.org/pdf/1704.07008v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"data-adaptive-statistics-for-multiple","repo_url":"https://github.com/wilsoncai1992/adaptest","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"astronomy","task_name":"Astronomy"},{"task_slug":"marketing","task_name":"Marketing"},{"task_slug":"hypothesis-testing","task_name":"Two-sample testing"},{"task_slug":"high","task_name":"Vocal Bursts Intensity Prediction"},{"task_slug":null,"task_name":"valid"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}