FlashAttention and Efficient Attention
The Problem with Standard Attention
FlashAttention
Key Insight: Don't materialize the full attention matrix.
Technique:
Results:
Other Optimizations
1. Multi-Query Attention (MQA)
2. Grouped-Query Attention (GQA)
3. Sliding Window Attention
Context Length Evolution
Model | Context Length |
|-------|---------------|
GPT-2 | 1K |
GPT-3 | 4K |
GPT-4 | 32K-128K |
Claude 3 | 200K |
Gemini 1.5 | 1M+ |