Apply a Fit/Transform Pipeline

~20 mincode completion

Implement run_pipeline(X_train, X_test, spec).

spec is a list of step names, applied in order. Two steps exist:

  • "center" subtracts the training column mean
  • "scale" divides by the training column standard deviation (population, ddof=0); if a column's training standard deviation is 0, divide by 1 instead

Fit every step on the training data, in order, using the training data as it stands after the previous steps. Then transform X_test through the same fitted steps and return the transformed test data.

An empty spec returns X_test unchanged.

Worked example:

X_train = [[1], [3], [5]]   X_test = [[7]]   spec = ["center"]

training  = 3
test output = [[7 - 3]] = [[4]]

Note what does not happen: the test mean is 7, so a step that refit on test data would return [[0]] and hide the shift entirely. Getting [[4]] is the point.

Hint: compute each step's parameters from the running training matrix, apply that step to both matrices, then move to the next step.

Examples

Single centering step: test row uses the training mean, not its own

Input
run_pipeline([[1], [3], [5]], [[7]], ["center"])
Output
[[4]]

Empty spec returns the test data unchanged

Input
run_pipeline([[1], [2]], [[5]], [])
Output
[[5]]

Center then scale, applied in order with training statistics

Input
run_pipeline([[0], [4]], [[2]], ["center", "scale"])
Output
[[0]]

Hints

Hint 1

picks between two values elementwise without branching.

Hint 2

Watch for this: refit each step on the test data leaking test statistics.

Requirements

  • X_train: (n_train, d) training matrix, used to fit every step

  • X_test: (n_test, d) test matrix, only ever transformed

  • spec: ordered list of step names, "center" and/or "scale"

  • Return (n_test, d) array of transformed test data.

Constraints

  • Allowed library: NumPy only

  • Time limit: 200 ms, Memory: 64 MB

Where this shows up

~20 min

8 employers weight this skill

3 AI product companies, 2 big tech firms, 2 enterprise vendors, 1 defense company. Top match scores 87.

Python
import numpy as np

def run_pipeline(X_train: np.ndarray, X_test: np.ndarray, spec: list) -> np.ndarray:
    """
    Fit a pipeline on X_train and return X_test transformed through it.

    Args:
        X_train: (n_train, d) training matrix, used to fit every step
        X_test:  (n_test, d) test matrix, only ever transformed
        spec:    ordered list of step names, "center" and/or "scale"

    Returns:
        (n_test, d) array of transformed test data.
    """
    # YOUR CODE HERE
    pass
Loading docs…

The AI Mentor needs an account

It reads your code and the failing tests and nudges you toward the fix without handing you the answer. Free accounts get it on every problem you're working on today.