{"url":"/method/pipedream-2bw","slug":"pipedream-2bw","name":"PipeDream-2BW","full_name":"PipeDream-2BW","full_name_withheld":false,"description_markdown":"**PipeDream-2BW** is an asynchronous pipeline parallel method that supports memory-efficient pipeline parallelism, a hybrid form of parallelism that combines data and model parallelism with input pipelining. PipeDream-2BW uses a novel pipelining and weight gradient coalescing strategy, combined with the double buffering of weights, to ensure high throughput, low memory footprint, and weight update semantics similar to data parallelism. In addition, PipeDream2BW automatically partitions the model over the available hardware resources, while respecting hardware constraints such as memory capacities of accelerators, and topologies and bandwidths of interconnects. PipeDream-2BW also determines when to employ existing memory-savings techniques, such as activation recomputation, that trade off extra computation for lower memory footprint.\r\n\r\nThe two main features are a double-buffered weight update (2BW) and flush mechanisms ensure high throughput. PipeDream-2BW\r\nsplits models into stages over multiple workers, and each stage is replicated an equal number of times (with data-parallel updates across replicas of the same stage).  Such parallel pipelines work well for models where each layer is repeated a fixed number of times (e.g., [transformer](https://paperswithcode.com/method/transformer) models).","description_state":"present","introduced_year":null,"introduced_by":{"title":"Memory-Efficient Pipeline-Parallel DNN Training","paper":"/paper/memory-efficient-pipeline-parallel-dnn","first_author":"Deepak Narayanan","n_authors":5,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/memory-efficient-pipeline-parallel-dnn"},"source":{"url":"https://arxiv.org/abs/2006.09503v3","title":"Memory-Efficient Pipeline-Parallel DNN Training","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Asynchronous Pipeline Parallel","url":"/methods/category/asynchronous-pipeline-parallel","pwc_aliases":[]},{"area":"General","area_id":"general","collection":"Model Parallel Methods","url":"/methods/category/model-parallel-methods","pwc_aliases":[]},{"area":"General","area_id":"general","collection":"Distributed Methods","url":"/methods/category/distributed-methods","pwc_aliases":[]}],"n_papers_tagged":3,"archive_num_papers":3,"papers_newest_first":[{"paper":"/paper/pipeoptim-ensuring-effective-1f1b-schedule","title":"PipeOptim: Ensuring Effective 1F1B Schedule with Optimizer-Dependent Weight Prediction","date":"2023-12-01","arxiv_id":"2312.00839","n_code_links":1,"syntology":null},{"paper":"/paper/group-based-interleaved-pipeline-parallelism","title":"Group-based Interleaved Pipeline Parallelism for Large-scale DNN Training","date":"2021-09-29","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/memory-efficient-pipeline-parallel-dnn","title":"Memory-Efficient Pipeline-Parallel DNN Training","date":"2020-06-16","arxiv_id":"2006.09503","n_code_links":1,"syntology":{"ran":2,"of":10,"unverified":8,"pointer_only":0}}],"papers_shown":3,"tasks":[{"task":"/task/image-classification","name":"Image Classification","papers":1},{"task":"/task/machine-translation","name":"Machine Translation","papers":1},{"task":"/task/prediction","name":"Prediction","papers":1},{"task":"/task/sentiment-analysis","name":"Sentiment Analysis","papers":1},{"task":"/task/image-classification","name":"image-classification","papers":1}],"tasks_shown":5,"n_tasks":5,"usage_by_year":[{"year":"2020","papers":1},{"year":"2021","papers":1},{"year":"2023","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/pipedream-2bw"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}