Leakage-Free Feature Standardization

~20 mincode completion

Implement scale_with_train_stats(X_train, X_test) that returns standardized X_test.

Examples

2-feature train range [0,2]: test point [1,3] → [0.0, 2.0]

Input
scale_with_train_stats([[0, 0], [2, 2]], [[1, 3]])
Output
[[0, 2]]

Train with unit std: test point shifted and scaled correctly

Input
scale_with_train_stats([[1, 2], [3, 4]], [[5, 6]])
Output
[[3, 3]]

Single feature: two test points standardized by train stats

Input
scale_with_train_stats([[2], [4], [6]], [[2], [8]])
Output
[[-1.22474], [2.44949]]

Hints

Hint 1

picks between two values elementwise without branching.

Hint 2

Watch for this: computed stats over combined train and test set.

Requirements

  • X_train: Training features, shape (n_train, d)

  • X_test: Test features, shape (n_test, d)

  • Return Standardized X_test of shape (n_test, d).

Constraints

  • Allowed library: NumPy only

  • Time limit: 200 ms, Memory: 64 MB

Where this shows up

~20 min

8 employers weight this skill

3 big tech firms, 2 AI product companies, 1 defense company, 1 enterprise vendor, 1 quant fund. Top match scores 85.

Python
import numpy as np

def scale_with_train_stats(X_train: np.ndarray, X_test: np.ndarray) -> np.ndarray:
    """
    Standardize X_test using mean and std computed from X_train only.

    Args:
        X_train: Training features, shape (n_train, d)
        X_test:  Test features, shape (n_test, d)

    Returns:
        Standardized X_test of shape (n_test, d).
    """
    # YOUR CODE HERE
    pass
Loading docs…

The AI Mentor needs an account

It reads your code and the failing tests and nudges you toward the fix without handing you the answer. Free accounts get it on every problem you're working on today.