Q-Learning Update

~25 mincode completion

Implement q_learning_update(Q, states, actions, rewards, next_states, dones, alpha, gamma).

  • Q has shape (n_states, n_actions).
  • The five transition arrays are parallel and 1D; dones uses 1 for terminal and 0 otherwise.
  • Return the updated Q-table. Do not modify the caller's array in place; copy it first.

Examples

Single non-terminal transition bootstraps off max Q(s')

Input
q_learning_update([[0, 0], [0, 10]], [0], [1], [1], [1], [0], 0.5, 0.9)
Output
[[0, 5], [0, 10]]

Terminal transition uses the reward alone, no bootstrap

Input
q_learning_update([[0, 0], [0, 10]], [0], [1], [1], [1], [1], 0.5, 0.9)
Output
[[0, 0.5], [0, 10]]

Sequential updates: the second transition sees the first one's write

Input
q_learning_update([[0, 0], [0, 0], [0, 0]], [2, 1, 0], [0, 0, 0], [1, 0, 0], [2, 2, 1], [1, 0, 0], 1, 0.9)
Output
[[0.81, 0], [0.9, 0], [1, 0]]

Hints

Hint 1

Walk the input once and accumulate as you go.

Hint 2

Watch for this: bootstrapped through terminal transitions ignoring the done flag.

Requirements

  • Q: (n_states, n_actions) Q-table

  • states: (n,) state indices

  • actions: (n,) action indices

  • rewards: (n,) rewards

  • next_states: (n,) next-state indices

  • dones: (n,) 1 if the transition ended the episode, else 0

  • alpha: learning rate

  • gamma: discount factor

  • Return Updated (n_states, n_actions) Q-table. The input Q is not mutated.

  • Use a fully vectorised implementation without Python loops

Constraints

  • Vectorised implementation only, no Python loops

  • Allowed library: NumPy only

  • Time limit: 200 ms, Memory: 64 MB

Where this shows up

~25 min

5 employers weight this skill

3 autonomy companies, 2 defense companies. Top match scores 85.

Python
import numpy as np

def q_learning_update(Q: np.ndarray, states: np.ndarray, actions: np.ndarray,
                      rewards: np.ndarray, next_states: np.ndarray,
                      dones: np.ndarray, alpha: float, gamma: float) -> np.ndarray:
    """
    Apply tabular Q-learning updates sequentially, one transition at a time.

    Args:
        Q:           (n_states, n_actions) Q-table
        states:      (n,) state indices
        actions:     (n,) action indices
        rewards:     (n,) rewards
        next_states: (n,) next-state indices
        dones:       (n,) 1 if the transition ended the episode, else 0
        alpha:       learning rate
        gamma:       discount factor

    Returns:
        Updated (n_states, n_actions) Q-table. The input Q is not mutated.
    """
    # YOUR CODE HERE
    pass

Run your code to see results

⌘↵ runs against the visible tests

Loading docs…

The AI Mentor needs an account

It reads your code and the failing tests and nudges you toward the fix without handing you the answer. Free accounts get it on every problem you're working on today.