Attention head
One of several parallel attention computations; more heads = richer attention, bigger KV cache.
One of several parallel attention computations; more heads = richer attention, bigger KV cache.
One of several parallel attention computations; more heads = richer attention, bigger KV cache.
One of several parallel attention computations; more heads = richer attention, bigger KV cache.