{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/understanding-synthetic-gradients-and","title":"Understanding Synthetic Gradients and Decoupled Neural Interfaces","arxiv_id":"1703.00522","date":"2017-03-01","proceeding":"ICML 2017 8","authors":["Wojciech Marian Czarnecki","Grzegorz Świrszcz","Max Jaderberg","Simon Osindero","Oriol Vinyals","Koray Kavukcuoglu"],"abstract":"When training neural networks, the use of Synthetic Gradients (SG) allows\nlayers or modules to be trained without update locking - without waiting for a\ntrue error gradient to be backpropagated - resulting in Decoupled Neural\nInterfaces (DNIs). This unlocked ability of being able to update parts of a\nneural network asynchronously and with only local information was demonstrated\nto work empirically in Jaderberg et al (2016). However, there has been very\nlittle demonstration of what changes DNIs and SGs impose from a functional,\nrepresentational, and learning dynamics point of view. In this paper, we study\nDNIs through the use of synthetic gradients on feed-forward networks to better\nunderstand their behaviour and elucidate their effect on optimisation. We show\nthat the incorporation of SGs does not affect the representational strength of\nthe learning system for a neural network, and prove the convergence of the\nlearning system for linear and deep linear models. On practical problems we\ninvestigate the mechanism by which synthetic gradient estimators approximate\nthe true loss, and, surprisingly, how that leads to drastically different\nlayer-wise representations. Finally, we also expose the relationship of using\nsynthetic gradients to other error approximation techniques and find a unifying\nlanguage for discussion and comparison.","url_abs":"http://arxiv.org/abs/1703.00522v1","url_pdf":"http://arxiv.org/pdf/1703.00522v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"understanding-synthetic-gradients-and","repo_url":"https://github.com/quangvu0702/Synthetic-Gradients","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1703.00522","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}