Skip to content
Navigation
Dashboard
🎬Video•20 min

Regularization: Dropout & BatchNorm

Master techniques to prevent overfitting and stabilize training.

Regularization Techniques

Dropout

Randomly "drop" neurons with probability p during training.

Benefits:

  • Prevents co-adaptation of neurons
  • Acts as ensemble of subnetworks
  • Batch Normalization

    Normalizes layer inputs to stabilize training.

    Steps: 1. Compute mean and variance of batch 2. Normalize: x̂ = (x - μ) / √(σ² + ε) 3. Scale and shift: y = γx̂ + β

    Benefits:

  • Allows higher learning rates
  • Faster convergence
  • Acts as slight regularization
  • Layer Normalization

    Used in Transformers instead of BatchNorm.

  • Works with variable sequence lengths
  • No dependency on batch statistics
  • 🎯 Key Takeaways

    • ✓Dropout prevents overfitting by randomly dropping neurons
    • ✓Batch Normalization stabilizes training
    • ✓Layer Normalization is preferred for Transformers
    • ✓Combine multiple regularization techniques

    📚 Additional Resources