Skip to content
Navigation
Dashboard
🎬Video•30 min

BERT and Bidirectional Models

Understand BERT and encoder-only language models.

BERT: Bidirectional Encoder Representations

What is BERT?

BERT = Bidirectional Encoder Representations from Transformers

Released by Google in 2018, it revolutionized NLP.

Pre-training Objectives

1. Masked Language Modeling (MLM)

  • Mask 15% of tokens randomly
  • Predict the masked tokens
  • Learns bidirectional context
  • 2. Next Sentence Prediction (NSP)

  • Predict if sentence B follows A
  • Helps with understanding relationships
  • BERT Architecture

  • 12 layers (BERT-base) / 24 layers (BERT-large)
  • 768/1024 hidden dimensions
  • 12/16 attention heads
  • 110M/340M parameters
  • Fine-tuning BERT

    BERT is pre-trained on large corpus, then fine-tuned for:

  • Text classification
  • Named Entity Recognition
  • Question Answering
  • Sentence similarity
  • BERT Variants

  • RoBERTa: Better training, no NSP
  • ALBERT: Parameter sharing for efficiency
  • DistilBERT: Smaller, faster
  • DeBERTa: Disentangled attention, better performance
  • 🎯 Key Takeaways

    • ✓BERT uses bidirectional attention for understanding
    • ✓MLM masks tokens for the model to predict
    • ✓Pre-train once, fine-tune for many tasks
    • ✓RoBERTa and DeBERTa are improved variants