Gradient Checkpointing

Appears in 1 paper · 1 tutorial

A memory optimisation technique where intermediate activations are discarded during forward pass and recomputed during backward pass.

As used in Paper 19 — Ring Attention with Blockwise Transformers for Near-Infinite Context →

A memory optimisation technique where intermediate activations are discarded during forward pass and recomputed during backward pass. Reduces memory at the cost of extra computation. Complements Ring Attention for training.

As used in Fine-Tuning & Model Customization →

Saving memory by recomputing intermediate activations during backprop instead of storing them; costs extra compute. (M08)