gpu-memory-utilization

Appears in 1 tutorial

Flag for the fraction of GPU memory vLLM may use; the key throughput/stability knob (bigger KV cache vs less headroom).

As used in LLM Infrastructure →

Flag for the fraction of GPU memory vLLM may use; the key throughput/stability knob (bigger KV cache vs less headroom).