{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-theory-for-distribution-regression","title":"Learning Theory for Distribution Regression","arxiv_id":"1411.2066","date":"2014-11-08","proceeding":null,"authors":["Zoltan Szabo","Bharath Sriperumbudur","Barnabas Poczos","Arthur Gretton"],"abstract":"We focus on the distribution regression problem: regressing to vector-valued\noutputs from probability measures. Many important machine learning and\nstatistical tasks fit into this framework, including multi-instance learning\nand point estimation problems without analytical solution (such as\nhyperparameter or entropy estimation). Despite the large number of available\nheuristics in the literature, the inherent two-stage sampled nature of the\nproblem makes the theoretical analysis quite challenging, since in practice\nonly samples from sampled distributions are observable, and the estimates have\nto rely on similarities computed between sets of points. To the best of our\nknowledge, the only existing technique with consistency guarantees for\ndistribution regression requires kernel density estimation as an intermediate\nstep (which often performs poorly in practice), and the domain of the\ndistributions to be compact Euclidean. In this paper, we study a simple,\nanalytically computable, ridge regression-based alternative to distribution\nregression, where we embed the distributions to a reproducing kernel Hilbert\nspace, and learn the regressor from the embeddings to the outputs. Our main\ncontribution is to prove that this scheme is consistent in the two-stage\nsampled setup under mild conditions (on separable topological domains enriched\nwith kernels): we present an exact computational-statistical efficiency\ntrade-off analysis showing that our estimator is able to match the one-stage\nsampled minimax optimal rate [Caponnetto and De Vito, 2007; Steinwart et al.,\n2009]. This result answers a 17-year-old open question, establishing the\nconsistency of the classical set kernel [Haussler, 1999; Gaertner et. al, 2002]\nin regression. We also cover consistency for more recent kernels on\ndistributions, including those due to [Christmann and Steinwart, 2010].","url_abs":"http://arxiv.org/abs/1411.2066v4","url_pdf":"http://arxiv.org/pdf/1411.2066v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-theory-for-distribution-regression","repo_url":"https://bitbucket.org/szzoli/ite","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"density-estimation","task_name":"Density Estimation"},{"task_slug":"learning-theory","task_name":"Learning Theory"},{"task_slug":"open-question","task_name":"Open-Ended Question Answering"},{"task_slug":"regression-1","task_name":"regression"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1411.2066","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}