Self-Attention Deep Dive
What is Self-Attention?
Each token attends to all tokens in the same sequence (including itself).
Step-by-Step Computation
1. Create Q, K, V: Q = X × W_Q K = X × W_K V = X × W_V
2. Compute Attention Scores: scores = Q × K^T
3. Scale: scaled_scores = scores / √d_k (Prevents softmax saturation)
4. Apply Softmax: attention_weights = softmax(scaled_scores)
5. Weighted Sum: output = attention_weights × V
Masking
Padding Mask: Ignore padding tokens Causal Mask: Prevent looking at future tokens (for generation)
Complexity
O(n²) for sequence length n