Methods › General › Replicated Data Parallel

Replicated Data Parallel

4 methods 6 papers tagged archive 2025-07-28

This section contains a compilation of distributed methods for scaling deep learning to very large models. There are many different strategies for scaling training across multiple devices, including:

Image credit: Jordi Torres.

Methods

All 4 methods in this collection, most-tagged first. Year is the archive's introduced_year; the archive stores 2000 when it has none, shown here as “–”. Papers counts distinct papers the archive tags with the method. Click a heading to sort.

PyTorch DDP – 3
BAGUA – 1
ByteScheduler 2019 1
DABMD Distributed Any-Batch Mirror Descent 2020 1