Skip to content
Navigation
Dashboard
🎬Video•30 min

Multi-Head Attention

Learn how multiple attention heads capture different relationships.

Multi-Head Attention

Why Multiple Heads?

Single attention head learns one "type" of relationship. Multiple heads can learn:

  • Syntactic relationships
  • Semantic relationships
  • Positional patterns
  • Different contextual features
  • How It Works

    1. Split: Divide Q, K, V into h heads 2. Attend: Each head computes attention independently 3. Concat: Concatenate all head outputs 4. Project: Linear transformation to original dimension

    MultiHead(Q, K, V) = Concat(head_1, ..., head_h) × W_O

    Typical Configurations

    Model | Heads | d_model | d_k |
    |-------|-------|---------|-----|
    BERT-base | 12 | 768 | 64 |
    GPT-2 | 12 | 768 | 64 |
    GPT-3 | 96 | 12288 | 128 |

    What Do Heads Learn?

    Research shows different heads specialize in:

  • Subject-verb agreement
  • Coreference resolution
  • Syntactic dependencies
  • Positional patterns
  • 🎯 Key Takeaways

    • ✓Multiple heads capture different types of relationships
    • ✓Each head operates on a projected subspace
    • ✓Outputs are concatenated and projected
    • ✓Different heads specialize in different patterns

    📚 Additional Resources