k-Nearest Neighbors Classifier

~20 minimplementation

Implement knn_predict(X_train, y_train, X_test, k).

  • X_train has shape (n_train, n_features), X_test has shape (n_test, n_features).
  • y_train is a 1D array of integer labels.
  • Return a 1D array of n_test predicted labels.

Hint: broadcasting gives you the whole distance matrix in one line: X_test[:, None, :] - X_train[None, :, :]. You never need to rank distances, but it costs nothing here.

Examples

k=3, two well-separated clusters: point near the origin cluster

Input
knn_predict([[0, 0], [1, 0], [0, 1], [5, 5], [6, 5]], [0, 0, 0, 1, 1], [[0.5, 0.5]], 3)
Output
[0]

k=1 reduces to nearest-neighbor lookup

Input
knn_predict([[0, 0], [1, 0], [0, 1], [5, 5], [6, 5]], [0, 0, 0, 1, 1], [[5.2, 5.1], [0.1, 0.2]], 1)
Output
[1, 0]

Larger k pulls in the majority class across the boundary

Input
knn_predict([[0, 0], [1, 0], [0, 1], [5, 5], [6, 5]], [0, 0, 0, 1, 1], [[4, 4]], 5)
Output
[0]

Hints

Hint 1

You need the index of the extreme value, not the value itself.

Hint 2

Watch for this: used squared distance but compared against a sqrt threshold.

Requirements

  • X_train: (n_train, n_features) training features

  • y_train: (n_train,) integer labels

  • X_test: (n_test, n_features) points to classify

  • k: number of neighbors to consider

  • Return (n_test,) array of predicted integer labels.

  • Use a fully vectorised implementation without Python loops

Constraints

  • Vectorised implementation only, no Python loops

  • Allowed library: NumPy only

  • Time limit: 200 ms, Memory: 64 MB

Where this shows up

~20 min

7 employers weight this skill

2 health and bio companies, 2 quant funds, 2 enterprise vendors, 1 AI product company. Top match scores 75.

Python
import numpy as np

def knn_predict(X_train: np.ndarray, y_train: np.ndarray,
                X_test: np.ndarray, k: int) -> np.ndarray:
    """
    Predict labels for X_test by majority vote over the k nearest
    training points under Euclidean distance.

    Ties in the vote are broken by choosing the smallest label.

    Args:
        X_train: (n_train, n_features) training features
        y_train: (n_train,) integer labels
        X_test:  (n_test, n_features) points to classify
        k:       number of neighbors to consider

    Returns:
        (n_test,) array of predicted integer labels.
    """
    # YOUR CODE HERE
    pass

Run your code to see results

⌘↵ runs against the visible tests

Loading docs…

The AI Mentor needs an account

It reads your code and the failing tests and nudges you toward the fix without handing you the answer. Free accounts get it on every problem you're working on today.