{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/variance-reduction-in-sgd-by-distributed","title":"Variance Reduction in SGD by Distributed Importance Sampling","arxiv_id":"1511.06481","date":"2015-11-20","proceeding":null,"authors":["Guillaume Alain","Alex Lamb","Chinnadhurai Sankar","Aaron Courville","Yoshua Bengio"],"abstract":"Humans are able to accelerate their learning by selecting training materials\nthat are the most informative and at the appropriate level of difficulty. We\npropose a framework for distributing deep learning in which one set of workers\nsearch for the most informative examples in parallel while a single worker\nupdates the model on examples selected by importance sampling. This leads the\nmodel to update using an unbiased estimate of the gradient which also has\nminimum variance when the sampling proposal is proportional to the L2-norm of\nthe gradient. We show experimentally that this method reduces gradient variance\neven in a context where the cost of synchronization across machines cannot be\nignored, and where the factors for importance sampling are not updated\ninstantly across the training set.","url_abs":"http://arxiv.org/abs/1511.06481v7","url_pdf":"http://arxiv.org/pdf/1511.06481v7.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"variance-reduction-in-sgd-by-distributed","repo_url":"https://github.com/idiap/importance-sampling","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1511.06481","atlas_url":"https://app.syntology.ai/?focus=1511.06481","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}