Skip to content
Navigation
Dashboard
💻Interactive•35 min

Self-Attention Mechanism

Implement and understand self-attention from scratch.

Self-Attention Deep Dive

What is Self-Attention?

Each token attends to all tokens in the same sequence (including itself).

Step-by-Step Computation

1. Create Q, K, V: Q = X × W_Q K = X × W_K V = X × W_V

2. Compute Attention Scores: scores = Q × K^T

3. Scale: scaled_scores = scores / √d_k (Prevents softmax saturation)

4. Apply Softmax: attention_weights = softmax(scaled_scores)

5. Weighted Sum: output = attention_weights × V

Masking

Padding Mask: Ignore padding tokens Causal Mask: Prevent looking at future tokens (for generation)

Complexity

O(n²) for sequence length n

  • Challenge for very long sequences
  • Led to innovations like FlashAttention
  • 🎯 Key Takeaways

    • ✓Self-attention relates each token to all others
    • ✓Scaling by √d_k prevents softmax saturation
    • ✓Masking controls what tokens can attend to
    • ✓O(n²) complexity limits sequence length

    📚 Additional Resources