The TD Error
Implement td_errors(rewards, values, next_values, dones, gamma) returning one per step.
Examples
Exactly as predicted: the error is zero and nothing is learned
- Input
- td_errors([0, 0], [1, 1], [1, 1], [0, 0], 1)
- Output
- [0, 0]
An unexpected reward is a positive surprise
- Input
- td_errors([1], [0], [0], [0], 1)
- Output
- [1]
A terminal step ignores the next value entirely, however large
- Input
- td_errors([10], [2], [99], [1], 0.99)
- Output
- [8]
Hints
Hint 1
Convert the input with before doing elementwise work.
Hint 2
A common slip here: bootstrapped through a terminal state instead of masking it.
Requirements
rewards: (T,) reward received at each step: (T,) V(s_t)
next_values: (T,) V(s_{t+1})dones: (T,) 1 if the episode ended at this step, else 0gamma: discount factorReturn (T,) array of TD errors
Constraints
Allowed library: NumPy only
Time limit: 200 ms, Memory: 64 MB
Try similar problems(4)
Reinforcement Learning: Rewards, Senses and PPO · ~14 min
Reinforcement Learning: Rewards, Senses and PPO · ~12 min
Reinforcement Learning: Rewards, Senses and PPO · ~14 min
Reinforcement Learning: Rewards, Senses and PPO · ~12 min
import numpy as np
def td_errors(rewards, values, next_values, dones, gamma):
"""
One-step temporal-difference error at every step of a batch.
Args:
rewards: (T,) reward received at each step
values: (T,) V(s_t)
next_values: (T,) V(s_{t+1})
dones: (T,) 1 if the episode ended at this step, else 0
gamma: discount factor
Returns:
(T,) array of TD errors
"""
# YOUR CODE HERE
pass