Skip to content
Navigation
Dashboard
🎬Video•25 min

Word Embeddings (Word2Vec, GloVe)

Understand how words are represented as dense vectors.

Word Embeddings

From One-Hot to Dense Vectors

One-Hot Encoding:

  • Sparse, high-dimensional
  • No semantic meaning
  • "cat" = [1,0,0,...], "dog" = [0,1,0,...]
  • Word Embeddings:

  • Dense, low-dimensional (50-300)
  • Capture semantic relationships
  • Similar words → similar vectors
  • Word2Vec

    Skip-gram: Predict context from word CBOW: Predict word from context

    Famous example: king - man + woman ≈ queen

    GloVe (Global Vectors)

  • Uses co-occurrence statistics
  • Combines local and global context
  • Often performs better on analogy tasks
  • Modern Embeddings

    Contextual Embeddings (BERT, GPT):

  • Same word gets different vectors based on context
  • "bank" (river) ≠ "bank" (financial)
  • Much more powerful than static embeddings
  • 🎯 Key Takeaways

    • ✓Embeddings represent words as dense vectors
    • ✓Word2Vec learns from local context windows
    • ✓GloVe uses global co-occurrence statistics
    • ✓Modern models use contextual embeddings

    📚 Additional Resources