{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/nonparametric-semi-supervised-learning-of","title":"Nonparametric semi-supervised learning of class proportions","arxiv_id":"1601.01944","date":"2016-01-08","proceeding":null,"authors":["Shantanu Jain","Martha White","Michael W. Trosset","Predrag Radivojac"],"abstract":"The problem of developing binary classifiers from positive and unlabeled data\nis often encountered in machine learning. A common requirement in this setting\nis to approximate posterior probabilities of positive and negative classes for\na previously unseen data point. This problem can be decomposed into two steps:\n(i) the development of accurate predictors that discriminate between positive\nand unlabeled data, and (ii) the accurate estimation of the prior probabilities\nof positive and negative examples. In this work we primarily focus on the\nlatter subproblem. We study nonparametric class prior estimation and formulate\nthis problem as an estimation of mixing proportions in two-component mixture\nmodels, given a sample from one of the components and another sample from the\nmixture itself. We show that estimation of mixing proportions is generally\nill-defined and propose a canonical form to obtain identifiability while\nmaintaining the flexibility to model any distribution. We use insights from\nthis theory to elucidate the optimization surface of the class priors and\npropose an algorithm for estimating them. To address the problems of\nhigh-dimensional density estimation, we provide practical transformations to\nlow-dimensional spaces that preserve class priors. Finally, we demonstrate the\nefficacy of our method on univariate and multivariate data.","url_abs":"http://arxiv.org/abs/1601.01944v1","url_pdf":"http://arxiv.org/pdf/1601.01944v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"nonparametric-semi-supervised-learning-of","repo_url":"https://github.com/dzeiberg/alphamax","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"density-estimation","task_name":"Density Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1601.01944","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}