Methods › General › Sharded Data Parallel Methods › ZeRO-Infinity
ZeRO-Infinity
Introduced by Samyam Rajbhandari et al. in ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
ZeRO-Infinity is a sharded data parallel system that extends ZeRO with new innovations in heterogeneous memory access called the infinity offload engine. This allows ZeRO-Infinity to support massive model sizes on limited GPU resources by exploiting CPU and NVMe memory simultaneously. In addition, ZeRO-Infinity also introduces a novel GPU memory optimization technique called memory-centric tiling to support extremely large individual layers that would otherwise not fit in GPU memory even one layer at a time.
Papers archive 2025-07-28
2 shown of 2, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage 6 Jun 2025 · 0 repositories · arXiv:2506.06472
-
ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning 16 Apr 2021 · 0 repositories · arXiv:2104.07857
Tasks archive 2025-07-28
3 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
| Task | Papers |
|---|---|
| CPU | 2 |
| GPU | 2 |
| Large Language Model | 1 |
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections