Multi-Head Attention
Why Multiple Heads?
Single attention head learns one "type" of relationship. Multiple heads can learn:
How It Works
1. Split: Divide Q, K, V into h heads 2. Attend: Each head computes attention independently 3. Concat: Concatenate all head outputs 4. Project: Linear transformation to original dimension
MultiHead(Q, K, V) = Concat(head_1, ..., head_h) × W_O
Typical Configurations
Model | Heads | d_model | d_k |
|-------|-------|---------|-----|
BERT-base | 12 | 768 | 64 |
GPT-2 | 12 | 768 | 64 |
GPT-3 | 96 | 12288 | 128 |
What Do Heads Learn?
Research shows different heads specialize in: