Linear Layer Weight Gradient

~15 mincode completion

Implement linear_weight_grad(X, d_out) that returns .

Examples

2x2 input, 2x1 upstream gradient

Input
linear_weight_grad([[1, 2], [3, 4]], [[1], [1]])
Output
[[4], [6]]

Identity input: gradient equals d_out

Input
linear_weight_grad([[1, 0], [0, 1]], [[2], [3]])
Output
[[2], [3]]

Single sample, 3 features, 1 output

Input
linear_weight_grad([[1, 2, 3]], [[1]])
Output
[[1], [2], [3]]

Hints

Hint 1

Use a matrix product rather than nested loops, and check which operand transposes.

Hint 2

Watch for this: used X dot d out without transpose.

Requirements

  • X: Input to the linear layer, shape (m, n)

  • d_out: Upstream gradient dL/dY, shape (m, k)

  • Return Weight gradient dL/dW, shape (n, k) = X.T @ d_out

Constraints

  • Allowed library: NumPy only

  • Time limit: 200 ms, Memory: 64 MB

Where this shows up

~15 min

8 employers weight this skill

4 frontier labs, 2 big tech firms, 1 autonomy company, 1 enterprise vendor. Top match scores 92.

Python
import numpy as np

def linear_weight_grad(X: np.ndarray, d_out: np.ndarray) -> np.ndarray:
    """
    Compute the gradient of the loss w.r.t. the weight matrix W
    of a linear layer Y = X @ W.

    Args:
        X:     Input to the linear layer, shape (m, n)
        d_out: Upstream gradient dL/dY, shape (m, k)

    Returns:
        Weight gradient dL/dW, shape (n, k) = X.T @ d_out
    """
    # YOUR CODE HERE
    pass
Loading docs…

The AI Mentor needs an account

It reads your code and the failing tests and nudges you toward the fix without handing you the answer. Free accounts get it on every problem you're working on today.