Fine-tuning Head Gradient
~20 mincode completion
Implement finetune_gradient(Z, W, y_true) that returns of shape (d, 1).
Examples
Perfect predictions: zero gradient
- Input
- finetune_gradient([[1, 0], [0, 1]], [[1], [1]], [[1], [1]])
- Output
- [[0], [0]]
Known gradient: overprediction pushes W down
- Input
- finetune_gradient([[1, 0], [0, 1]], [[2], [2]], [[1], [1]])
- Output
- [[1], [1]]
Hints
Hint 1
Use a matrix product rather than nested loops, and check which operand transposes.
Hint 2
Do not forget to factor 2 over m. That step is easy to skip.
Requirements
Z: Frozen embeddings, shape (m, d): Head weight vector, shape (d, 1)
y_true: Targets, shape (m, 1)Return Gradient dL/dW of shape (d, 1).
Constraints
Allowed library: NumPy only
Time limit: 200 ms, Memory: 64 MB
Where this shows up
~20 min
••••••••••••••••
8 employers weight this skill
4 frontier labs, 3 AI product companies, 1 enterprise vendor. Top match scores 78.
Python
import numpy as np
def finetune_gradient(Z: np.ndarray, W: np.ndarray, y_true: np.ndarray) -> np.ndarray:
"""
Compute gradient of MSE loss w.r.t. the linear head weights W.
Args:
Z: Frozen embeddings, shape (m, d)
W: Head weight vector, shape (d, 1)
y_true: Targets, shape (m, 1)
Returns:
Gradient dL/dW of shape (d, 1).
"""
# YOUR CODE HERE
pass