{"url":"/method/gpipe","slug":"gpipe","name":"GPipe","full_name":"GPipe","full_name_withheld":false,"description_markdown":"**GPipe** is a distributed model parallel method for neural networks. With GPipe, each model can be specified as a sequence of layers, and consecutive groups of layers can be partitioned into cells. Each cell is then placed on a separate accelerator. Based on this partitioned setup, batch splitting is applied. A mini-batch of training examples is split into smaller micro-batches, then the execution of each set of micro-batches is pipelined over cells. Synchronous mini-batch gradient descent is applied for training, where gradients are accumulated across all micro-batches in a mini-batch and applied at the end of a mini-batch.","description_state":"present","introduced_year":null,"introduced_by":{"title":"GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism","paper":"/paper/gpipe-efficient-training-of-giant-neural","first_author":"Yanping Huang","n_authors":11,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/gpipe-efficient-training-of-giant-neural"},"source":{"url":"https://arxiv.org/abs/1811.06965v5","title":"GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/kakaobrain/torchgpipe/blob/a1b4ee25574864e7650e7905a69ce156da9752ec/torchgpipe/gpipe.py#L134","code_snippet_url_on_a_code_host":true,"categories":[{"area":"General","area_id":"general","collection":"Synchronous Pipeline Parallel","url":"/methods/category/synchronous-pipeline-parallel","pwc_aliases":[]},{"area":"General","area_id":"general","collection":"Model Parallel Methods","url":"/methods/category/model-parallel-methods","pwc_aliases":[]},{"area":"General","area_id":"general","collection":"Distributed Methods","url":"/methods/category/distributed-methods","pwc_aliases":[]}],"n_papers_tagged":7,"archive_num_papers":7,"papers_newest_first":[{"paper":"/paper/pipeoptim-ensuring-effective-1f1b-schedule","title":"PipeOptim: Ensuring Effective 1F1B Schedule with Optimizer-Dependent Weight Prediction","date":"2023-12-01","arxiv_id":"2312.00839","n_code_links":1,"syntology":null},{"paper":"/paper/hydra-a-system-for-large-multi-model-deep","title":"Hydra: A System for Large Multi-Model Deep Learning","date":"2021-10-16","arxiv_id":"2110.08633","n_code_links":1,"syntology":null},{"paper":"/paper/group-based-interleaved-pipeline-parallelism","title":"Group-based Interleaved Pipeline Parallelism for Large-scale DNN Training","date":"2021-09-29","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":null,"title":"Automatic Graph Partitioning for Very Large-scale Deep Learning","date":"2021-03-30","arxiv_id":"2103.16063","n_code_links":0,"syntology":null},{"paper":null,"title":"Analyzing the Performance of Graph Neural Networks with Pipe Parallelism","date":"2020-12-20","arxiv_id":"2012.10840","n_code_links":0,"syntology":null},{"paper":"/paper/torchgpipe-on-the-fly-pipeline-parallelism","title":"torchgpipe: On-the-fly Pipeline Parallelism for Training Giant Models","date":"2020-04-21","arxiv_id":"2004.09910","n_code_links":3,"syntology":null},{"paper":"/paper/gpipe-efficient-training-of-giant-neural","title":"GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism","date":"2018-11-16","arxiv_id":"1811.06965","n_code_links":13,"syntology":{"ran":1,"of":25,"unverified":24,"pointer_only":16}}],"papers_shown":7,"tasks":[{"task":"/task/deep-learning","name":"Deep Learning","papers":3},{"task":"/task/image-classification","name":"Image Classification","papers":2},{"task":"/task/machine-translation","name":"Machine Translation","papers":2},{"task":"/task/image-classification","name":"image-classification","papers":2},{"task":"/task/machine-learning","name":"BIG-bench Machine Learning","papers":1},{"task":"/task/edge-classification","name":"Edge Classification","papers":1},{"task":"/task/fine-grained-image-classification","name":"Fine-Grained Image Classification","papers":1},{"task":null,"name":"GPU","papers":1},{"task":"/task/language-modeling","name":"Language Modeling","papers":1},{"task":"/task/language-modelling","name":"Language Modelling","papers":1},{"task":"/task/link-prediction","name":"Link Prediction","papers":1},{"task":"/task/model-selection","name":"Model Selection","papers":1},{"task":"/task/prediction","name":"Prediction","papers":1},{"task":"/task/protein-folding","name":"Protein Folding","papers":1},{"task":"/task/scheduling","name":"Scheduling","papers":1},{"task":"/task/sentiment-analysis","name":"Sentiment Analysis","papers":1},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":1},{"task":"/task/translation","name":"Translation","papers":1},{"task":"/task/graph-partitioning","name":"graph partitioning","papers":1}],"tasks_shown":19,"n_tasks":19,"usage_by_year":[{"year":"2018","papers":1},{"year":"2020","papers":2},{"year":"2021","papers":3},{"year":"2023","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/gpipe"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}