Skip to content
Navigation
Dashboard
📖Reading•25 min

Positional Encodings

Understand how transformers capture sequence order.

Positional Encodings

The Problem

Attention is permutation-invariant: "cat sat" = "sat cat" We need to inject position information.

Sinusoidal Positional Encoding (Original)

PE(pos, 2i) = sin(pos / 10000^(2i/d)) PE(pos, 2i+1) = cos(pos / 10000^(2i/d))

Properties:

  • Unique encoding for each position
  • Relative positions can be computed
  • Generalizes to unseen lengths
  • Learned Positional Embeddings

    Used by: BERT, GPT-2, GPT-3

  • Learn position embeddings during training
  • Limited to max sequence length
  • Often work better in practice
  • Relative Positional Encoding

    Used by: T5, Transformer-XL

  • Encode relative distance, not absolute position
  • Better for variable-length sequences
  • Rotary Position Embedding (RoPE)

    Used by: LLaMA, GPT-NeoX

  • Encodes position in the rotation of Q and K
  • Good for long sequences
  • State-of-the-art for many models
  • 🎯 Key Takeaways

    • ✓Transformers need explicit position information
    • ✓Sinusoidal encodings generalize to any length
    • ✓Learned embeddings often work better
    • ✓RoPE is state-of-the-art for modern LLMs

    📚 Additional Resources