Pipelines
~12 mincode completion
Write build_and_score(X_train, y_train, X_test, y_test) that builds a plus pipeline, fits it on the training data, and returns its accuracy on the test data as a float.
Pass random_state=0 to so the result is reproducible.
Examples
A cleanly separable problem is classified perfectly
- Input
- build_and_score([[0], [1], [2], [8], [9], [10]], [0, 0, 0, 1, 1, 1], [[0.5], [9.5]], [0, 1])
- Output
- 1
Two features, still separable, still perfect
- Input
- build_and_score([[0, 0], [1, 1], [9, 9], [10, 10]], [0, 0, 1, 1], [[0.5, 0.5], [9.5, 9.5]], [0, 1])
- Output
- 1
Hints
Hint 1
Work directly with the arguments X_train, y_train, X_test, y_test and return the result rather than printing it.
Hint 2
Watch for this: scaled outside the pipeline.
Requirements
Return Accuracy on the test set, as a float.
Constraints
Allowed library: NumPy only
Time limit: 200 ms, Memory: 64 MB
Try similar problems(4)
Where this shows up
~12 min
••••••••••••••
7 employers weight this skill
2 health and bio companies, 2 quant funds, 2 enterprise vendors, 1 AI product company. Top match scores 43.
Python
import numpy as np
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
def build_and_score(X_train, y_train, X_test, y_test):
"""
Build a scaler + classifier pipeline, fit it, and score it.
Args:
X_train, y_train: training data
X_test, y_test: held-out data
Returns:
Accuracy on the test set, as a float.
"""
# YOUR CODE HERE
pass