The chain rule made computational. Learn how gradients flow backward through a network to train every parameter simultaneously.
MSE Loss Gradient
~15 min· Hard
Sigmoid Gradient (Backprop)
ReLU Backward Pass
~10 min· Easy
Linear Layer Weight Gradient
~15 min· Medium
Sign in for the concept check
Optional multiple-choice questions on the ideas behind this section. Most useful after you have tried the coding problems above.