Scaling Without Leaking
Write scale_train_test(X_train, X_test) that fits a on the training data and returns the scaled test data.
The test data is deliberately chosen so that a leaking implementation, which refits on the test set, gives a different answer from a correct one.
Examples
Test data is scaled with the training mean and std, so it does not come out centred on zero
- Input
- scale_train_test([[0], [2], [4]], [[10], [12]])
- Output
- [[4.89898], [6.12372]]
A test point equal to the training mean scales to exactly 0
- Input
- scale_train_test([[0], [10]], [[5]])
- Output
- [[0]]
Hints
Hint 1
Work directly with the arguments X_train, X_test and return the result rather than printing it.
Hint 2
Watch for this: fit on test data.
Requirements
X_train: 2-D training featuresX_test: 2-D test featuresReturn the scaled test features.
Constraints
Allowed library: NumPy only
Time limit: 200 ms, Memory: 64 MB
Where this shows up
7 employers weight this skill
2 health and bio companies, 2 quant funds, 2 enterprise vendors, 1 AI product company. Top match scores 43.
import numpy as np
from sklearn.preprocessing import StandardScaler
def scale_train_test(X_train, X_test):
"""
Scale the test set using statistics learned from the training set.
Args:
X_train: 2-D training features
X_test: 2-D test features
Returns:
The scaled test features.
"""
# YOUR CODE HERE
pass