{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/neuralfdr-learning-discovery-thresholds-from","title":"NeuralFDR: Learning Discovery Thresholds from Hypothesis Features","arxiv_id":"1711.01312","date":"2017-11-03","proceeding":"NeurIPS 2017 12","authors":["Fei Xia","Martin J. Zhang","James Zou","David Tse"],"abstract":"As datasets grow richer, an important challenge is to leverage the full\nfeatures in the data to maximize the number of useful discoveries while\ncontrolling for false positives. We address this problem in the context of\nmultiple hypotheses testing, where for each hypothesis, we observe a p-value\nalong with a set of features specific to that hypothesis. For example, in\ngenetic association studies, each hypothesis tests the correlation between a\nvariant and the trait. We have a rich set of features for each variant (e.g.\nits location, conservation, epigenetics etc.) which could inform how likely the\nvariant is to have a true association. However popular testing approaches, such\nas Benjamini-Hochberg's procedure (BH) and independent hypothesis weighting\n(IHW), either ignore these features or assume that the features are categorical\nor uni-variate. We propose a new algorithm, NeuralFDR, which automatically\nlearns a discovery threshold as a function of all the hypothesis features. We\nparametrize the discovery threshold as a neural network, which enables flexible\nhandling of multi-dimensional discrete and continuous features as well as\nefficient end-to-end optimization. We prove that NeuralFDR has strong false\ndiscovery rate (FDR) guarantees, and show that it makes substantially more\ndiscoveries in synthetic and real datasets. Moreover, we demonstrate that the\nlearned discovery threshold is directly interpretable.","url_abs":"http://arxiv.org/abs/1711.01312v4","url_pdf":"http://arxiv.org/pdf/1711.01312v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"neuralfdr-learning-discovery-thresholds-from","repo_url":"https://github.com/fxia22/NeuralFDR","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}