Normalize activations within a mini-batch to keep training stable. BatchNorm lets you use higher learning rates and reduces sensitivity to initialization.
Batch Normalization Forward
~15 min· Medium
Layer Normalization
~12 min· Easy
Sign in for the concept check
Optional multiple-choice questions on the ideas behind this section. Most useful after you have tried the coding problems above.