Skip to content
Navigation
Dashboard
🎬Video•35 min

GPT Series (GPT-2, GPT-3, GPT-4)

Trace the evolution of OpenAI's GPT models.

The GPT Series

GPT-1 (2018)

  • 117M parameters
  • Proved unsupervised pre-training works
  • Fine-tuned for downstream tasks
  • GPT-2 (2019)

  • 1.5B parameters
  • "Too dangerous to release" (initially)
  • Zero-shot capabilities emerged
  • Trained on WebText (40GB)
  • GPT-3 (2020)

  • 175B parameters
  • Few-shot learning without fine-tuning
  • In-context learning
  • API-based access (ChatGPT foundation)
  • GPT-4 (2023)

  • Multimodal (text + images)
  • Much better reasoning
  • 100K+ context length (GPT-4 Turbo)
  • Powers ChatGPT Plus
  • Key Innovations

    Scaling Laws:

  • More parameters → Better performance
  • More data → Better performance
  • Predictable improvement with scale
  • Emergent Capabilities:

  • Abilities that appear at scale
  • Not present in smaller models
  • Chain-of-thought reasoning
  • 🎯 Key Takeaways

    • ✓GPT models use decoder-only architecture
    • ✓Scaling leads to emergent capabilities
    • ✓GPT-3 introduced few-shot learning
    • ✓GPT-4 added multimodal understanding

    📚 Additional Resources