{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/gradient-estimation-using-stochastic","title":"Gradient Estimation Using Stochastic Computation Graphs","arxiv_id":"1506.05254","date":"2015-06-17","proceeding":"NeurIPS 2015 12","authors":["John Schulman","Nicolas Heess","Theophane Weber","Pieter Abbeel"],"abstract":"In a variety of problems originating in supervised, unsupervised, and\nreinforcement learning, the loss function is defined by an expectation over a\ncollection of random variables, which might be part of a probabilistic model or\nthe external world. Estimating the gradient of this loss function, using\nsamples, lies at the core of gradient-based learning algorithms for these\nproblems. We introduce the formalism of stochastic computation\ngraphs---directed acyclic graphs that include both deterministic functions and\nconditional probability distributions---and describe how to easily and\nautomatically derive an unbiased estimator of the loss function's gradient. The\nresulting algorithm for computing the gradient estimator is a simple\nmodification of the standard backpropagation algorithm. The generic scheme we\npropose unifies estimators derived in variety of prior work, along with\nvariance-reduction techniques therein. It could assist researchers in\ndeveloping intricate models involving a combination of stochastic and\ndeterministic operations, enabling, for example, attention, memory, and control\nactions.","url_abs":"http://arxiv.org/abs/1506.05254v3","url_pdf":"http://arxiv.org/pdf/1506.05254v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"gradient-estimation-using-stochastic","repo_url":"https://github.com/caravagn/GDA","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1506.05254","atlas_url":"https://app.syntology.ai/?focus=1506.05254","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}