Attention Head
An independent parallel compute sub-unit inside self-attention layers that tracks relationships between words.
Last reviewed: July 25, 2026
An attention head is one of several parallel, independent sub-units within a transformer’s self-attention layer, each learning to focus on different types of relationships between tokens in a sequence. Multi-head attention — running several attention heads in parallel rather than one large one — is a defining architectural feature of the transformer, introduced in the original 2017 “Attention Is All You Need” paper.
What Each Head Does
Every attention head independently computes its own Query, Key, and Value projections of the input token representations, then produces its own attention pattern — a weighted combination of other tokens’ information, weighted by how relevant each token is to the current one. Different heads in a trained transformer tend to specialize in recognizing different kinds of relationships: some heads have been observed to track syntactic structure (like subject-verb agreement), others track long-range coreference (linking a pronoun back to the noun it refers to), and others attend mostly to nearby tokens for local context.
Why Multiple Heads Instead of One
A single attention mechanism computing one weighted average per token can only represent one “type” of relationship at a time for a given token. Splitting attention into multiple smaller heads that each operate on a lower-dimensional slice of the representation lets the model simultaneously track several different kinds of relationships in parallel, then combine their outputs — this turns out to produce noticeably better performance than a single larger attention mechanism with the same total parameter count.
Practical Relevance
The number of attention heads is a core architectural hyperparameter set when a model is designed (for example, Llama 3 70B uses 64 attention heads), and it directly affects the size of the KV cache during inference — architectural variants like multi-query attention and grouped-query attention modify how many heads share key-value projections specifically to reduce this memory cost.
Attention Head Pruning
Because different heads specialize to different degrees and some contribute more to model output than others, research has found that a meaningful fraction of a trained transformer’s attention heads can often be pruned entirely with minimal accuracy loss — evidence that not every head learns an equally important or unique function, and that models are often somewhat over-provisioned with more heads than strictly necessary for their task performance. This finding underlies some structured pruning techniques that specifically target redundant attention heads as a compression strategy, though identifying which heads are safe to remove typically requires empirical evaluation on the target model and task rather than being predictable from the architecture alone.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.