Skip to content
Navigation
Dashboard
📖Reading•20 min

Layer Normalization & Residuals

Understand the components that make deep transformers trainable.

Layer Normalization & Residual Connections

Residual Connections (Skip Connections)

output = x + SubLayer(x)

Benefits:

  • Enable training very deep networks
  • Gradients flow directly through
  • Help with vanishing gradients
  • Layer Normalization

    Normalize across features (not batch): LayerNorm(x) = γ × (x - μ) / σ + β

    Pre-LN vs Post-LN:

  • Post-LN (original): LayerNorm after sublayer
  • Pre-LN (modern): LayerNorm before sublayer
  • Pre-LN is more stable for deep networks
  • Transformer Block

    Pre-LN Transformer Block: 1. x = x + Attention(LayerNorm(x)) 2. x = x + FFN(LayerNorm(x))

    Feed-Forward Network

    FFN(x) = GELU(x × W₁ + b₁) × W₂ + b₂

  • Usually 4× hidden dimension
  • Applies independently to each position
  • 🎯 Key Takeaways

    • ✓Residual connections enable deep networks
    • ✓Layer Normalization stabilizes training
    • ✓Pre-LN is preferred for modern transformers
    • ✓FFN expands and contracts the representation

    📚 Additional Resources