Perceptrons and Activation Functions
The Perceptron
The fundamental building block of neural networks:
output = activation(Σ(wᵢ × xᵢ) + b)
Why Activation Functions?
Without them, neural networks are just linear transformations. Activation functions introduce non-linearity.
Common Activation Functions
1. Sigmoid - σ(x) = 1 / (1 + e⁻ˣ)
2. ReLU - ReLU(x) = max(0, x)
3. GELU - Used in Transformers (BERT, GPT)
4. Softmax - Output: Probability distribution