Methods › General › Stochastic Optimization › Gradient Checkpointing
Gradient Checkpointing
Introduced by Tianqi Chen et al. in Training Deep Nets with Sublinear Memory Cost
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Gradient Checkpointing is a method used for reducing the memory footprint when training deep neural networks, at the cost of having a small increase in computation time.
Papers archive 2025-07-28
14 shown of 14, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Optimal Gradient Checkpointing for Sparse and Recurrent Architectures using Off-Chip Memory 16 Dec 2024 · 0 repositories · arXiv:2412.11810
-
Look Every Frame All at Once: Video-Ma²mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing 29 Nov 2024 · 0 repositories · arXiv:2411.19460
-
Superior Scoring Rules for Probabilistic Evaluation of Single-Label Multi-Class Classification Tasks 25 Jul 2024 · 1 repository · arXiv:2407.17697
-
A Study of Optimizations for Fine-tuning Large Language Models 4 Jun 2024 · 0 repositories · arXiv:2406.02290
-
DITTO: Diffusion Inference-Time T-Optimization for Music Generation 22 Jan 2024 · 0 repositories · arXiv:2401.12179
-
CAPIVARA: Cost-Efficient Approach for Improving Multilingual CLIP Performance on Low-Resource Languages 20 Oct 2023 · 1 repository · arXiv:2310.13683
-
Unsupervised Discovery of Interpretable Directions in h-space of Pre-trained Diffusion Models 15 Oct 2023 · 0 repositories · arXiv:2310.09912
-
DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training 5 Oct 2023 · 1 repository · arXiv:2310.03294
-
Colossal-Auto: Unified Automation of Parallelization and Activation Checkpoint for Large-scale Models 6 Feb 2023 · 1 repository · arXiv:2302.02599
-
GLEAM: Greedy Learning for Large-Scale Accelerated MRI Reconstruction 18 Jul 2022 · 1 repository · arXiv:2207.08393
-
Combined Scaling for Zero-shot Transfer Learning 19 Nov 2021 · 0 repositories · arXiv:2111.10050
-
Doc2Dict: Information Extraction as Text Generation 16 May 2021 · 1 repository · arXiv:2105.07510
-
Self-supervised Pretraining of Visual Features in the Wild 2 Mar 2021 · 1 repository · arXiv:2103.01988
-
Training Deep Nets with Sublinear Memory Cost 21 Apr 2016 · 6 repositories · arXiv:1604.06174Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)
Tasks archive 2025-07-28
20 shown of 29 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections