Q-Learning Update
Implement q_learning_update(Q, states, actions, rewards, next_states, dones, alpha, gamma).
Qhas shape(n_states, n_actions).- The five transition arrays are parallel and 1D;
donesuses1for terminal and0otherwise. - Return the updated Q-table. Do not modify the caller's array in place; copy it first.
Examples
Single non-terminal transition bootstraps off max Q(s')
- Input
- q_learning_update([[0, 0], [0, 10]], [0], [1], [1], [1], [0], 0.5, 0.9)
- Output
- [[0, 5], [0, 10]]
Terminal transition uses the reward alone, no bootstrap
- Input
- q_learning_update([[0, 0], [0, 10]], [0], [1], [1], [1], [1], 0.5, 0.9)
- Output
- [[0, 0.5], [0, 10]]
Sequential updates: the second transition sees the first one's write
- Input
- q_learning_update([[0, 0], [0, 0], [0, 0]], [2, 1, 0], [0, 0, 0], [1, 0, 0], [2, 2, 1], [1, 0, 0], 1, 0.9)
- Output
- [[0.81, 0], [0.9, 0], [1, 0]]
Hints
Hint 1
Walk the input once and accumulate as you go.
Hint 2
Watch for this: bootstrapped through terminal transitions ignoring the done flag.
Requirements
Q: (n_states, n_actions) Q-tablestates: (n,) state indicesactions: (n,) action indicesrewards: (n,) rewardsnext_states: (n,) next-state indicesdones: (n,) 1 if the transition ended the episode, else 0alpha: learning rategamma: discount factorReturn Updated (n_states, n_actions) Q-table. The input Q is not mutated.
Use a fully vectorised implementation without Python loops
Constraints
Vectorised implementation only, no Python loops
Allowed library: NumPy only
Time limit: 200 ms, Memory: 64 MB
Where this shows up
5 employers weight this skill
3 autonomy companies, 2 defense companies. Top match scores 85.
import numpy as np
def q_learning_update(Q: np.ndarray, states: np.ndarray, actions: np.ndarray,
rewards: np.ndarray, next_states: np.ndarray,
dones: np.ndarray, alpha: float, gamma: float) -> np.ndarray:
"""
Apply tabular Q-learning updates sequentially, one transition at a time.
Args:
Q: (n_states, n_actions) Q-table
states: (n,) state indices
actions: (n,) action indices
rewards: (n,) rewards
next_states: (n,) next-state indices
dones: (n,) 1 if the transition ended the episode, else 0
alpha: learning rate
gamma: discount factor
Returns:
Updated (n_states, n_actions) Q-table. The input Q is not mutated.
"""
# YOUR CODE HERE
pass