{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/on-learning-from-label-proportions","title":"On Learning from Label Proportions","arxiv_id":"1402.5902","date":"2014-02-24","proceeding":null,"authors":["Felix X. Yu","Krzysztof Choromanski","Sanjiv Kumar","Tony Jebara","Shih-Fu Chang"],"abstract":"Learning from Label Proportions (LLP) is a learning setting, where the\ntraining data is provided in groups, or \"bags\", and only the proportion of each\nclass in each bag is known. The task is to learn a model to predict the class\nlabels of the individual instances. LLP has broad applications in political\nscience, marketing, healthcare, and computer vision. This work answers the\nfundamental question, when and why LLP is possible, by introducing a general\nframework, Empirical Proportion Risk Minimization (EPRM). EPRM learns an\ninstance label classifier to match the given label proportions on the training\ndata. Our result is based on a two-step analysis. First, we provide a VC bound\non the generalization error of the bag proportions. We show that the bag sample\ncomplexity is only mildly sensitive to the bag size. Second, we show that under\nsome mild assumptions, good bag proportion prediction guarantees good instance\nlabel prediction. The results together provide a formal guarantee that the\nindividual labels can indeed be learned in the LLP setting. We discuss\napplications of the analysis, including justification of LLP algorithms,\nlearning with population proportions, and a paradigm for learning algorithms\nwith privacy guarantees. We also demonstrate the feasibility of LLP based on a\ncase study in real-world setting: predicting income based on census data.","url_abs":"http://arxiv.org/abs/1402.5902v2","url_pdf":"http://arxiv.org/pdf/1402.5902v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"on-learning-from-label-proportions","repo_url":"https://github.com/felixyu/pSVM","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"marketing","task_name":"Marketing"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1402.5902","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}