Scaling Without Leaking

~12 mincode completion

Write scale_train_test(X_train, X_test) that fits a on the training data and returns the scaled test data.

The test data is deliberately chosen so that a leaking implementation, which refits on the test set, gives a different answer from a correct one.

Examples

Test data is scaled with the training mean and std, so it does not come out centred on zero

Input
scale_train_test([[0], [2], [4]], [[10], [12]])
Output
[[4.89898], [6.12372]]

A test point equal to the training mean scales to exactly 0

Input
scale_train_test([[0], [10]], [[5]])
Output
[[0]]

Hints

Hint 1

Work directly with the arguments X_train, X_test and return the result rather than printing it.

Hint 2

Watch for this: fit on test data.

Requirements

  • X_train: 2-D training features

  • X_test: 2-D test features

  • Return the scaled test features.

Constraints

  • Allowed library: NumPy only

  • Time limit: 200 ms, Memory: 64 MB

Where this shows up

~12 min

7 employers weight this skill

2 health and bio companies, 2 quant funds, 2 enterprise vendors, 1 AI product company. Top match scores 43.

Python
import numpy as np
from sklearn.preprocessing import StandardScaler

def scale_train_test(X_train, X_test):
    """
    Scale the test set using statistics learned from the training set.

    Args:
        X_train: 2-D training features
        X_test: 2-D test features

    Returns:
        The scaled test features.
    """
    # YOUR CODE HERE
    pass

Run your code to see results

⌘↵ runs against the visible tests

Loading docs…

The AI Mentor needs an account

It reads your code and the failing tests and nudges you toward the fix without handing you the answer. Free accounts get it on every problem you're working on today.