{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/iterative-hard-thresholding-for-model","title":"Iterative Hard Thresholding for Model Selection in Genome-Wide Association Studies","arxiv_id":"1608.01398","date":"2016-08-04","proceeding":null,"authors":["Kevin L. Keys","Gary K. Chen","Kenneth Lange"],"abstract":"A genome-wide association study (GWAS) correlates marker variation with trait\nvariation in a sample of individuals. Each study subject is genotyped at a\nmultitude of SNPs (single nucleotide polymorphisms) spanning the genome. Here\nwe assume that subjects are unrelated and collected at random and that trait\nvalues are normally distributed or transformed to normality. Over the past\ndecade, researchers have been remarkably successful in applying GWAS analysis\nto hundreds of traits. The massive amount of data produced in these studies\npresent unique computational challenges. Penalized regression with LASSO or MCP\npenalties is capable of selecting a handful of associated SNPs from millions of\npotential SNPs. Unfortunately, model selection can be corrupted by false\npositives and false negatives, obscuring the genetic underpinning of a trait.\nThis paper introduces the iterative hard thresholding (IHT) algorithm to the\nGWAS analysis of continuous traits. Our parallel implementation of IHT\naccommodates SNP genotype compression and exploits multiple CPU cores and\ngraphics processing units (GPUs). This allows statistical geneticists to\nleverage commodity desktop computers in GWAS analysis and to avoid\nsupercomputing. We evaluate IHT performance on both simulated and real GWAS\ndata and conclude that it reduces false positive and false negative rates while\nremaining competitive in computational time with penalized regression. Source\ncode is freely available at https://github.com/klkeys/IHT.jl.","url_abs":"http://arxiv.org/abs/1608.01398v3","url_pdf":"http://arxiv.org/pdf/1608.01398v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"iterative-hard-thresholding-for-model","repo_url":"https://github.com/klkeys/IHT.jl","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":null,"task_name":"CPU"},{"task_slug":"model-selection","task_name":"Model Selection"},{"task_slug":"regression-1","task_name":"regression"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}