Skip to content
Navigation
Dashboard
🎬Video•25 min

FlashAttention & Optimizations

Learn modern techniques for efficient attention computation.

FlashAttention and Efficient Attention

The Problem with Standard Attention

  • O(n²) memory for attention matrix
  • Memory bottleneck limits context length
  • GPU memory is expensive
  • FlashAttention

    Key Insight: Don't materialize the full attention matrix.

    Technique:

  • Tile-based computation
  • Fused operations
  • Recomputation instead of storage
  • Results:

  • 2-4x faster
  • 5-20x less memory
  • Enables longer contexts
  • Other Optimizations

    1. Multi-Query Attention (MQA)

  • Share K, V across heads
  • Reduces memory bandwidth
  • 2. Grouped-Query Attention (GQA)

  • Share K, V within groups of heads
  • Balance between MQA and full attention
  • Used in LLaMA 2, Mistral
  • 3. Sliding Window Attention

  • Attend to local window + global tokens
  • Used in Longformer, BigBird
  • Context Length Evolution

    Model | Context Length |
    |-------|---------------|
    GPT-2 | 1K |
    GPT-3 | 4K |
    GPT-4 | 32K-128K |
    Claude 3 | 200K |
    Gemini 1.5 | 1M+ |

    🎯 Key Takeaways

    • ✓FlashAttention avoids materializing full attention matrix
    • ✓Grouped-Query Attention balances speed and quality
    • ✓Modern LLMs support 100K+ context lengths
    • ✓Memory optimization enables longer contexts