Apply a Fit/Transform Pipeline
Implement run_pipeline(X_train, X_test, spec).
spec is a list of step names, applied in order. Two steps exist:
"center"subtracts the training column mean"scale"divides by the training column standard deviation (population,ddof=0); if a column's training standard deviation is0, divide by1instead
Fit every step on the training data, in order, using the training data as it stands after the previous steps. Then transform X_test through the same fitted steps and return the transformed test data.
An empty spec returns X_test unchanged.
Worked example:
X_train = [[1], [3], [5]] X_test = [[7]] spec = ["center"] training = 3 test output = [[7 - 3]] = [[4]]
Note what does not happen: the test mean is 7, so a step that refit on test data would return [[0]] and hide the shift entirely. Getting [[4]] is the point.
Hint: compute each step's parameters from the running training matrix, apply that step to both matrices, then move to the next step.
Examples
Single centering step: test row uses the training mean, not its own
- Input
- run_pipeline([[1], [3], [5]], [[7]], ["center"])
- Output
- [[4]]
Empty spec returns the test data unchanged
- Input
- run_pipeline([[1], [2]], [[5]], [])
- Output
- [[5]]
Center then scale, applied in order with training statistics
- Input
- run_pipeline([[0], [4]], [[2]], ["center", "scale"])
- Output
- [[0]]
Hints
Hint 1
picks between two values elementwise without branching.
Hint 2
Watch for this: refit each step on the test data leaking test statistics.
Requirements
X_train: (n_train, d) training matrix, used to fit every stepX_test: (n_test, d) test matrix, only ever transformedspec: ordered list of step names, "center" and/or "scale"Return (n_test, d) array of transformed test data.
Constraints
Allowed library: NumPy only
Time limit: 200 ms, Memory: 64 MB
Where this shows up
8 employers weight this skill
3 AI product companies, 2 big tech firms, 2 enterprise vendors, 1 defense company. Top match scores 87.
import numpy as np
def run_pipeline(X_train: np.ndarray, X_test: np.ndarray, spec: list) -> np.ndarray:
"""
Fit a pipeline on X_train and return X_test transformed through it.
Args:
X_train: (n_train, d) training matrix, used to fit every step
X_test: (n_test, d) test matrix, only ever transformed
spec: ordered list of step names, "center" and/or "scale"
Returns:
(n_test, d) array of transformed test data.
"""
# YOUR CODE HERE
pass