{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/on-the-optimality-of-averaging-in-distributed","title":"On the Optimality of Averaging in Distributed Statistical Learning","arxiv_id":"1407.2724","date":"2014-07-10","proceeding":null,"authors":["Jonathan Rosenblatt","Boaz Nadler"],"abstract":"A common approach to statistical learning with big-data is to randomly split\nit among $m$ machines and learn the parameter of interest by averaging the $m$\nindividual estimates. In this paper, focusing on empirical risk minimization,\nor equivalently M-estimation, we study the statistical error incurred by this\nstrategy. We consider two large-sample settings: First, a classical setting\nwhere the number of parameters $p$ is fixed, and the number of samples per\nmachine $n\\to\\infty$. Second, a high-dimensional regime where both\n$p,n\\to\\infty$ with $p/n \\to \\kappa \\in (0,1)$. For both regimes and under\nsuitable assumptions, we present asymptotically exact expressions for this\nestimation error. In the fixed-$p$ setting, under suitable assumptions, we\nprove that to leading order averaging is as accurate as the centralized\nsolution. We also derive the second order error terms, and show that these can\nbe non-negligible, notably for non-linear models. The high-dimensional setting,\nin contrast, exhibits a qualitatively different behavior: data splitting incurs\na first-order accuracy loss, which to leading order increases linearly with the\nnumber of machines. The dependence of our error approximations on the number of\nmachines traces an interesting accuracy-complexity tradeoff, allowing the\npractitioner an informed choice on the number of machines to deploy. Finally,\nwe confirm our theoretical analysis with several simulations.","url_abs":"http://arxiv.org/abs/1407.2724v2","url_pdf":"http://arxiv.org/pdf/1407.2724v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"on-the-optimality-of-averaging-in-distributed","repo_url":"https://github.com/johnros/ParalSimulate","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1407.2724","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}