From "what to look at" to the engine behind GPT and BERT. Scaled dot-product attention replaces recurrence with direct token-to-token relationships.
Softmax Attention Weights
~12 min· Medium
Scaled Dot-Product Attention
~25 min· Hard
Sinusoidal Positional Encoding
Causal Attention Mask
~6 min· Medium
Multi-Head Attention
~18 min· Hard
Query-Key Attention Scores
~12 min· Easy
Attention Context Vector
~10 min· Easy
Causal Multi-Head Self-Attention
~40 min· Expert
Decode One Token With a KV Cache
~20 min· Medium
Sign in for the concept check
Optional multiple-choice questions on the ideas behind this section. Most useful after you have tried the coding problems above.