Attention Is All You Need
The 2017 Revolution
Vaswani et al. introduced the Transformer, replacing recurrence entirely with attention.
Core Idea
Instead of processing sequentially, attend to all positions simultaneously:
Attention Mechanism
Attention(Q, K, V) = softmax(QK^T / √d_k) × V
Query (Q): What am I looking for? Key (K): What do I contain? Value (V): What information do I provide?
Transformer Architecture
Encoder:
Decoder:
Why It Works
1. Parallel computation: All positions at once 2. Direct connections: No information degradation 3. Learnable attention: Model decides relevance 4. Scalable: More compute = better performance