Methods › General › Sharded Data Parallel Methods › ZeRO
ZeRO
Introduced by Samyam Rajbhandari et al. in ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Zero Redundancy Optimizer (ZeRO) is a sharded data parallel method for distributed training. ZeRODP removes the memory state redundancies across data-parallel processes by partitioning the model states instead of replicating them, and it retains the compute/communication efficiency by retaining the computational granularity and communication volume of DP using a dynamic communication schedule during training.
Papers archive 2025-07-28
9 shown of 9, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Memory Analysis on the Training Course of DeepSeek Models 11 Feb 2025 · 0 repositories · arXiv:2502.07846
-
Accelerating Large Language Model Training with Hybrid GPU-based Compression 4 Sep 2024 · 0 repositories · arXiv:2409.02423
-
A Study of Optimizations for Fine-tuning Large Language Models 4 Jun 2024 · 0 repositories · arXiv:2406.02290
-
Zero redundancy distributed learning with differential privacy 20 Nov 2023 · 0 repositories · arXiv:2311.11822
-
Dissecting the Runtime Performance of the Training, Fine-tuning, and Inference of Large Language Models 7 Nov 2023 · 0 repositories · arXiv:2311.03687
-
ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models 16 Oct 2023 · 3 repositories · arXiv:2310.10505Syntology ran 8 of 13 samples · 5 unverified · 13 pointer-only (licence)
-
Rethinking Memory and Communication Cost for Efficient Large Language Model Training 9 Oct 2023 · 0 repositories · arXiv:2310.06003
-
ZeRO++: Extremely Efficient Collective Communication for Giant Model Training 16 Jun 2023 · 1 repository · arXiv:2306.10209
-
ZeRO: Memory Optimizations Toward Training Trillion Parameter Models 4 Oct 2019 · 10 repositories · arXiv:1910.02054Syntology ran 1 of 1 samples · 0 unverified
Tasks archive 2025-07-28
11 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections