FlashAttention

Appears in 1 tutorial

An exact, memory-efficient attention kernel that avoids writing big intermediates to slow memory; speeds attention, especially for long contexts.

As used in LLM Infrastructure →

An exact, memory-efficient attention kernel that avoids writing big intermediates to slow memory; speeds attention, especially for long contexts. No quality loss.