Splitting Train and Test
~10 mincode completion
Write split_sizes(X, y, test_size) that splits the data and returns [n_train, n_test], the number of rows in each half.
Use random_state=0 so the result is the same every time.
Examples
Holding back 20 percent of 10 rows leaves 8 for training
- Input
- split_sizes([[1], [2], [3], [4], [5], [6], [7], [8], [9], [10]], [0, 1, 0, 1, 0, 1, 0, 1, 0, 1], 0.2)
- Output
- [8, 2]
Holding back half of 10 rows splits it evenly
- Input
- split_sizes([[1], [2], [3], [4], [5], [6], [7], [8], [9], [10]], [0, 1, 0, 1, 0, 1, 0, 1, 0, 1], 0.5)
- Output
- [5, 5]
Hints
Hint 1
Work directly with the arguments X, y, test_size and return the result rather than printing it.
Hint 2
Watch for this: unpacked return values in wrong order.
Requirements
X: a 2-D feature arrayy: a 1-D label arraytest_size: the fraction to hold back, e.g. 0.2Return a list [n_train, n_test].
Constraints
Allowed library: NumPy only
Time limit: 200 ms, Memory: 64 MB
Where this shows up
~10 min
••••••••••••••
7 employers weight this skill
2 health and bio companies, 2 quant funds, 2 enterprise vendors, 1 AI product company. Top match scores 50.
Python
import numpy as np
from sklearn.model_selection import train_test_split
def split_sizes(X, y, test_size):
"""
Split the data and report how big each half is.
Args:
X: a 2-D feature array
y: a 1-D label array
test_size: the fraction to hold back, e.g. 0.2
Returns:
A list [n_train, n_test].
"""
# YOUR CODE HERE
pass