Leakage-Free Feature Standardization
~20 mincode completion
Implement scale_with_train_stats(X_train, X_test) that returns standardized X_test.
Examples
2-feature train range [0,2]: test point [1,3] → [0.0, 2.0]
- Input
- scale_with_train_stats([[0, 0], [2, 2]], [[1, 3]])
- Output
- [[0, 2]]
Train with unit std: test point shifted and scaled correctly
- Input
- scale_with_train_stats([[1, 2], [3, 4]], [[5, 6]])
- Output
- [[3, 3]]
Single feature: two test points standardized by train stats
- Input
- scale_with_train_stats([[2], [4], [6]], [[2], [8]])
- Output
- [[-1.22474], [2.44949]]
Hints
Hint 1
picks between two values elementwise without branching.
Hint 2
Watch for this: computed stats over combined train and test set.
Requirements
X_train: Training features, shape (n_train, d)X_test: Test features, shape (n_test, d)Return Standardized X_test of shape (n_test, d).
Constraints
Allowed library: NumPy only
Time limit: 200 ms, Memory: 64 MB
Where this shows up
~20 min
••••••••••••••••
8 employers weight this skill
3 big tech firms, 2 AI product companies, 1 defense company, 1 enterprise vendor, 1 quant fund. Top match scores 85.
Python
import numpy as np
def scale_with_train_stats(X_train: np.ndarray, X_test: np.ndarray) -> np.ndarray:
"""
Standardize X_test using mean and std computed from X_train only.
Args:
X_train: Training features, shape (n_train, d)
X_test: Test features, shape (n_test, d)
Returns:
Standardized X_test of shape (n_test, d).
"""
# YOUR CODE HERE
pass