{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-communication-efficient-distributed-3","title":"A Communication-Efficient Distributed Gradient Clipping Algorithm for Training Deep Neural Networks","arxiv_id":"2205.05040","date":"2022-05-10","proceeding":null,"authors":["Mingrui Liu","Zhenxun Zhuang","Yunwei Lei","Chunyang Liao"],"abstract":"In distributed training of deep neural networks, people usually run Stochastic Gradient Descent (SGD) or its variants on each machine and communicate with other machines periodically. However, SGD might converge slowly in training some deep neural networks (e.g., RNN, LSTM) because of the exploding gradient issue. Gradient clipping is usually employed to address this issue in the single machine setting, but exploring this technique in the distributed setting is still in its infancy: it remains mysterious whether the gradient clipping scheme can take advantage of multiple machines to enjoy parallel speedup. The main technical difficulty lies in dealing with nonconvex loss function, non-Lipschitz continuous gradient, and skipping communication rounds simultaneously. In this paper, we explore a relaxed-smoothness assumption of the loss landscape which LSTM was shown to satisfy in previous works, and design a communication-efficient gradient clipping algorithm. This algorithm can be run on multiple machines, where each machine employs a gradient clipping scheme and communicate with other machines after multiple steps of gradient-based updates. Our algorithm is proved to have $O\\left(\\frac{1}{N\\epsilon^4}\\right)$ iteration complexity and $O(\\frac{1}{\\epsilon^3})$ communication complexity for finding an $\\epsilon$-stationary point in the homogeneous data setting, where $N$ is the number of machines. This indicates that our algorithm enjoys linear speedup and reduced communication rounds. Our proof relies on novel analysis techniques of estimating truncated random variables, which we believe are of independent interest. Our experiments on several benchmark datasets and various scenarios demonstrate that our algorithm indeed exhibits fast convergence speed in practice and thus validates our theory.","url_abs":"https://arxiv.org/abs/2205.05040v2","url_pdf":"https://arxiv.org/pdf/2205.05040v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-communication-efficient-distributed-3","repo_url":"https://github.com/mingruiliu-ml-lab/communication-efficient-local-gradient-clipping","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"federated-learning","task_name":"Federated Learning"}],"methods":[{"method_slug":"gradient-clipping","method_name":"Gradient Clipping"},{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sgd","method_name":"SGD"},{"method_slug":"speed","method_name":"SPEED"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2205.05040","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2205.05040"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mingruiliu-ml-lab/communication-efficient-local-gradient-clipping","reach":null}],"summary":{"ran_draft_wrong":1,"ran_fixture":1,"ran_honours":1},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"9efb556e868426ed","entry":"batchify","repo":"mingruiliu-ml-lab/communication-efficient-local-gradient-clipping","repo_kind":"official","path":"nlp/main_lstm.py","file_url":"https://github.com/mingruiliu-ml-lab/communication-efficient-local-gradient-clipping/blob/HEAD/nlp/main_lstm.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"GPL-3.0","inline_ok":false,"mcp_get_code":{"code_sha256":"9efb556e868426ed"}},{"code_sha256_prefix":"2a0640f3091a0588","entry":"get_batch","repo":"mingruiliu-ml-lab/communication-efficient-local-gradient-clipping","repo_kind":"official","path":"nlp/main_lstm.py","file_url":"https://github.com/mingruiliu-ml-lab/communication-efficient-local-gradient-clipping/blob/HEAD/nlp/main_lstm.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"GPL-3.0","inline_ok":false,"mcp_get_code":{"code_sha256":"2a0640f3091a0588"}},{"code_sha256_prefix":"4372bb4533fd1936","entry":"repackage_hidden","repo":"mingruiliu-ml-lab/communication-efficient-local-gradient-clipping","repo_kind":"official","path":"nlp/main_lstm.py","file_url":"https://github.com/mingruiliu-ml-lab/communication-efficient-local-gradient-clipping/blob/HEAD/nlp/main_lstm.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"GPL-3.0","inline_ok":false,"mcp_get_code":{"code_sha256":"4372bb4533fd1936"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}