{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/gradient-surgery-for-multi-task-learning-1","title":"Gradient Surgery for Multi-Task Learning","arxiv_id":"2001.06782","date":"2020-01-19","proceeding":"NeurIPS 2020 12","authors":["Tianhe Yu","Saurabh Kumar","Abhishek Gupta","Sergey Levine","Karol Hausman","Chelsea Finn"],"abstract":"While deep learning and deep reinforcement learning (RL) systems have demonstrated impressive results in domains such as image classification, game playing, and robotic control, data efficiency remains a major challenge. Multi-task learning has emerged as a promising approach for sharing structure across multiple tasks to enable more efficient learning. However, the multi-task setting presents a number of optimization challenges, making it difficult to realize large efficiency gains compared to learning tasks independently. The reasons why multi-task learning is so challenging compared to single-task learning are not fully understood. In this work, we identify a set of three conditions of the multi-task optimization landscape that cause detrimental gradient interference, and develop a simple yet general approach for avoiding such interference between task gradients. We propose a form of gradient surgery that projects a task's gradient onto the normal plane of the gradient of any other task that has a conflicting gradient. On a series of challenging multi-task supervised and multi-task RL problems, this approach leads to substantial gains in efficiency and performance. Further, it is model-agnostic and can be combined with previously-proposed multi-task architectures for enhanced performance.","url_abs":"https://arxiv.org/abs/2001.06782v4","url_pdf":"https://arxiv.org/pdf/2001.06782v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"gradient-surgery-for-multi-task-learning-1","repo_url":"https://github.com/tianheyu927/PCGrad","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"gradient-surgery-for-multi-task-learning-1","repo_url":"https://github.com/OrthoDex/PCGrad-PyTorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"gradient-surgery-for-multi-task-learning-1","repo_url":"https://github.com/WeiChengTseng/Pytorch-PCGrad","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"gradient-surgery-for-multi-task-learning-1","repo_url":"https://github.com/avivnavon/nash-mtl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"gradient-surgery-for-multi-task-learning-1","repo_url":"https://github.com/cranial-xix/famo","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"gradient-surgery-for-multi-task-learning-1","repo_url":"https://github.com/grtzsohalf/SpeechNet-codebase","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"gradient-surgery-for-multi-task-learning-1","repo_url":"https://github.com/rangwani-harsh/PC_Grad_Pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"gradient-surgery-for-multi-task-learning-1","repo_url":"https://github.com/torchjd/torchjd","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"gradient-surgery-for-multi-task-learning-1","repo_url":"https://github.com/1hb6s7t/PCGrad-mindspore-example","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":{"status":"ok"}},{"paper_slug":"gradient-surgery-for-multi-task-learning-1","repo_url":"https://github.com/MindSpore-scientific-2/code-4/tree/main/PCGrad-mindspore-example","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"gradient-surgery-for-multi-task-learning-1","repo_url":"https://github.com/MindSpore-scientific-2/code-5/tree/main/PCGrad-mindspore-example","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"gradient-surgery-for-multi-task-learning-1","repo_url":"https://github.com/MindSpore-scientific-2/code-8/tree/main/PCGrad-mindspore-example","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"gradient-surgery-for-multi-task-learning-1","repo_url":"https://github.com/MindSpore-scientific/code-10/tree/main/PCGrad-mindspore-example","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"gradient-surgery-for-multi-task-learning-1","repo_url":"https://github.com/MindSpore-scientific/code-13/tree/main/PCGrad-mindspore-example","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"gradient-surgery-for-multi-task-learning-1","repo_url":"https://github.com/PaddlePaddle/PaddleScience","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"paddle","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"gradient-surgery-for-multi-task-learning-1","repo_url":"https://github.com/pwc-1/Paper-9/tree/main/4/PCGrad-mindspore-example","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"gradient-surgery-for-multi-task-learning-1","repo_url":"https://github.com/pwc-1/Paper-9/tree/main/7/PCGrad-mindspore-example","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"gradient-surgery-for-multi-task-learning-1","repo_url":"https://github.com/wgchang/PCGrad-pytorch-example","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"multi-task-learning","task_name":"Multi-Task Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2001.06782","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2001.06782"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/cranial-xix/famo","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindSpore-scientific/code-10/tree/main/PCGrad-mindspore-example","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/avivnavon/nash-mtl","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindSpore-scientific-2/code-8/tree/main/PCGrad-mindspore-example","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/1hb6s7t/PCGrad-mindspore-example","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/wgchang/PCGrad-pytorch-example","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/WeiChengTseng/Pytorch-PCGrad","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/PaddlePaddle/PaddleScience","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindSpore-scientific/code-13/tree/main/PCGrad-mindspore-example","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindSpore-scientific-2/code-5/tree/main/PCGrad-mindspore-example","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/grtzsohalf/SpeechNet-codebase","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindSpore-scientific-2/code-4/tree/main/PCGrad-mindspore-example","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/torchjd/torchjd","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/tianheyu927/PCGrad","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/OrthoDex/PCGrad-PyTorch","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/pwc-1/Paper-9/tree/main/7/PCGrad-mindspore-example","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/pwc-1/Paper-9/tree/main/4/PCGrad-mindspore-example","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/rangwani-harsh/PC_Grad_Pytorch","reach":null}],"summary":{"ran":6,"ran_draft_wrong":1,"unverified":3},"by_repo_kind":{"official":{"samples":1,"ran":0,"repositories":1},"listed":{"samples":9,"ran":7,"repositories":7}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"16a092a4692eb73d","entry":"PCGrad","repo":"grtzsohalf/SpeechNet-codebase","repo_kind":"listed","path":"src/pcgrad.py","file_url":"https://github.com/grtzsohalf/SpeechNet-codebase/blob/HEAD/src/pcgrad.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"16a092a4692eb73d"}},{"code_sha256_prefix":"5e1ac35b4889d3d9","entry":"PCGrad","repo":"WeiChengTseng/Pytorch-PCGrad","repo_kind":"listed","path":"pcgrad.py","file_url":"https://github.com/WeiChengTseng/Pytorch-PCGrad/blob/HEAD/pcgrad.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"5e1ac35b4889d3d9"}},{"code_sha256_prefix":"31aa3418224baf69","entry":"PCGrad","repo":"cranial-xix/famo","repo_kind":"listed","path":"methods/weight_methods.py","file_url":"https://github.com/cranial-xix/famo/blob/HEAD/methods/weight_methods.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"31aa3418224baf69"}},{"code_sha256_prefix":"c34a09ae701d0f40","entry":"PCGrad","repo":"avivnavon/nash-mtl","repo_kind":"listed","path":"methods/weight_methods.py","file_url":"https://github.com/avivnavon/nash-mtl/blob/HEAD/methods/weight_methods.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c34a09ae701d0f40"}},{"code_sha256_prefix":"654cdad96608431e","entry":"WeightMethod","repo":"cranial-xix/famo","repo_kind":"listed","path":"methods/weight_methods.py","file_url":"https://github.com/cranial-xix/famo/blob/HEAD/methods/weight_methods.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"654cdad96608431e"}},{"code_sha256_prefix":"9b136fce78a21563","entry":"WeightMethod","repo":"avivnavon/nash-mtl","repo_kind":"listed","path":"methods/weight_methods.py","file_url":"https://github.com/avivnavon/nash-mtl/blob/HEAD/methods/weight_methods.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"9b136fce78a21563"}},{"code_sha256_prefix":"4cf729e96cf2c9fe","entry":"pc_grad_update","repo":"OrthoDex/PCGrad-PyTorch","repo_kind":"listed","path":"pcgrad.py","file_url":"https://github.com/OrthoDex/PCGrad-PyTorch/blob/HEAD/pcgrad.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"4cf729e96cf2c9fe"}},{"code_sha256_prefix":"a64df2e1bec39655","entry":"PCGrad","repo":"tianheyu927/PCGrad","repo_kind":"official","path":"PCGrad_tf.py","file_url":"https://github.com/tianheyu927/PCGrad/blob/HEAD/PCGrad_tf.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a64df2e1bec39655"}},{"code_sha256_prefix":"0b13eb7dd654266c","entry":"PCGrad_backward","repo":"wgchang/PCGrad-pytorch-example","repo_kind":"listed","path":"pcgrad-example.py","file_url":"https://github.com/wgchang/PCGrad-pytorch-example/blob/HEAD/pcgrad-example.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"0b13eb7dd654266c"}},{"code_sha256_prefix":"3e75577924697c22","entry":"PCGrad_loss","repo":"rangwani-harsh/PC_Grad_Pytorch","repo_kind":"listed","path":"pc_grad_pytorch.py","file_url":"https://github.com/rangwani-harsh/PC_Grad_Pytorch/blob/HEAD/pc_grad_pytorch.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"3e75577924697c22"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}