Latency

Appears in 1 paper · 2 tutorials

The time taken to produce a response.

As used in Paper 23 — Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Model Parameters →

The time taken to produce a response. Higher N or more refinement rounds increase latency, which may be unacceptable in real-time applications even if accuracy improves.

As used in AI Production Engineering →

How long a response takes. Measured at percentiles; perceived latency (via streaming) matters as much as actual. (Mod 9)

As used in LLM Infrastructure →

How long one request takes (a per-user measure).