Gradient Boosting with Regression Stumps
Implement gradient_boost(X, y, n_rounds, lr) returning the final predictions FM on the training points as a 1D array of length n.
Xis a 1D array of one feature.yis a 1D array of targets.- is the learning rate ; small with many rounds is what makes boosting resistant to overfitting.
Watch the shrinkage. With lr < 1 the ensemble only closes a fraction of the gap each round. That is the point, not a bug.
Examples
One round, lr=1.0: the stump fits a perfect step function
- Input
- gradient_boost([1, 2, 3, 4], [1, 1, 5, 5], 1, 1)
- Output
- [1, 1, 5, 5]
Zero rounds returns the constant baseline mean(y)
- Input
- gradient_boost([1, 2, 3, 4], [1, 1, 5, 5], 0, 1)
- Output
- [3, 3, 3, 3]
Shrinkage: lr=0.5 closes only half the gap in one round
- Input
- gradient_boost([1, 2, 3, 4], [1, 1, 5, 5], 1, 0.5)
- Output
- [2, 2, 4, 4]
Hints
Hint 1
picks between two values elementwise without branching.
Hint 2
A common slip here: initialised F0 to zeros instead of mean y.
Requirements
X: (n,) 1D feature arrayy: (n,) target arrayn_rounds: number of boosting rounds: learning rate (shrinkage)
Return (n,) array of final predictions on the training points.
Constraints
Allowed library: NumPy only
Time limit: 200 ms, Memory: 64 MB
Where this shows up
8 employers weight this skill
2 big tech firms, 2 health and bio companies, 2 quant funds, 1 AI product company, 1 enterprise vendor. Top match scores 91.
import numpy as np
def gradient_boost(X: np.ndarray, y: np.ndarray, n_rounds: int, lr: float) -> np.ndarray:
"""
Gradient boosting with depth-1 regression stumps under squared loss.
Start from F0 = mean(y). Each round, fit a stump to the current residuals
and add lr * stump to the running prediction.
Split selection: candidate thresholds are midpoints between consecutive
values of np.unique(X); pick the threshold minimising total squared error,
breaking ties toward the smallest threshold.
Args:
X: (n,) 1D feature array
y: (n,) target array
n_rounds: number of boosting rounds
lr: learning rate (shrinkage)
Returns:
(n,) array of final predictions on the training points.
"""
# YOUR CODE HERE
pass