{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/staleness-aware-async-sgd-for-distributed","title":"Staleness-aware Async-SGD for Distributed Deep Learning","arxiv_id":"1511.05950","date":"2015-11-18","proceeding":null,"authors":["Wei Zhang","Suyog Gupta","Xiangru Lian","Ji Liu"],"abstract":"Deep neural networks have been shown to achieve state-of-the-art performance\nin several machine learning tasks. Stochastic Gradient Descent (SGD) is the\npreferred optimization algorithm for training these networks and asynchronous\nSGD (ASGD) has been widely adopted for accelerating the training of large-scale\ndeep networks in a distributed computing environment. However, in practice it\nis quite challenging to tune the training hyperparameters (such as learning\nrate) when using ASGD so as achieve convergence and linear speedup, since the\nstability of the optimization algorithm is strongly influenced by the\nasynchronous nature of parameter updates. In this paper, we propose a variant\nof the ASGD algorithm in which the learning rate is modulated according to the\ngradient staleness and provide theoretical guarantees for convergence of this\nalgorithm. Experimental verification is performed on commonly-used image\nclassification benchmarks: CIFAR10 and Imagenet to demonstrate the superior\neffectiveness of the proposed approach, compared to SSGD (Synchronous SGD) and\nthe conventional ASGD algorithm.","url_abs":"http://arxiv.org/abs/1511.05950v5","url_pdf":"http://arxiv.org/pdf/1511.05950v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"staleness-aware-async-sgd-for-distributed","repo_url":"https://github.com/Farhad-n/MultiGPU_Study","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"deep-learning","task_name":"Deep Learning"},{"task_slug":"distributed-computing","task_name":"Distributed Computing"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1511.05950","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1511.05950"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Farhad-n/MultiGPU_Study","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":1},"by_repo_kind":{"listed":{"samples":1,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"1baf5120601f6a0c","entry":"make_layers","repo":"Farhad-n/MultiGPU_Study","repo_kind":"listed","path":"vgg_fn.py","file_url":"https://github.com/Farhad-n/MultiGPU_Study/blob/HEAD/vgg_fn.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"1baf5120601f6a0c"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}