Skip to content
Navigation
Dashboard
💻Interactive•30 min

Tokenization Methods (BPE, WordPiece)

Master the tokenization methods that power modern language models.

Tokenization Methods

Why Subword Tokenization?

  • Word-level: Large vocabulary, OOV problem
  • Character-level: Long sequences, no word meaning
  • Subword: Best of both worlds!
  • Byte-Pair Encoding (BPE)

    Used by: GPT-2, GPT-3, LLaMA

    Algorithm: 1. Start with character vocabulary 2. Find most frequent pair 3. Merge into new token 4. Repeat until vocabulary size reached

    Example: "lowest" → ["low", "est"]

    WordPiece

    Used by: BERT, DistilBERT

    Similar to BPE but uses likelihood instead of frequency. Uses "##" prefix for continuation: "playing" → ["play", "##ing"]

    SentencePiece

  • Language-agnostic
  • Works directly on raw text
  • Used by: T5, LLaMA
  • Tokenizer Comparison

    Model | Tokenizer | Vocab Size |
    |-------|-----------|------------|
    GPT-4 | BPE | ~100K |
    BERT | WordPiece | 30K |
    LLaMA | SentencePiece | 32K |

    🎯 Key Takeaways

    • ✓Subword tokenization balances vocabulary and sequence length
    • ✓BPE merges frequent character pairs iteratively
    • ✓WordPiece is used by BERT models
    • ✓Tokenization directly impacts model performance

    📚 Additional Resources