The TD Error

~14 mincode completion

Implement td_errors(rewards, values, next_values, dones, gamma) returning one per step.

Examples

Exactly as predicted: the error is zero and nothing is learned

Input
td_errors([0, 0], [1, 1], [1, 1], [0, 0], 1)
Output
[0, 0]

An unexpected reward is a positive surprise

Input
td_errors([1], [0], [0], [0], 1)
Output
[1]

A terminal step ignores the next value entirely, however large

Input
td_errors([10], [2], [99], [1], 0.99)
Output
[8]

Hints

Hint 1

Convert the input with before doing elementwise work.

Hint 2

A common slip here: bootstrapped through a terminal state instead of masking it.

Requirements

  • rewards: (T,) reward received at each step

  • : (T,) V(s_t)

  • next_values: (T,) V(s_{t+1})

  • dones: (T,) 1 if the episode ended at this step, else 0

  • gamma: discount factor

  • Return (T,) array of TD errors

Constraints

  • Allowed library: NumPy only

  • Time limit: 200 ms, Memory: 64 MB

Python
import numpy as np


def td_errors(rewards, values, next_values, dones, gamma):
    """
    One-step temporal-difference error at every step of a batch.

    Args:
        rewards:     (T,) reward received at each step
        values:      (T,) V(s_t)
        next_values: (T,) V(s_{t+1})
        dones:       (T,) 1 if the episode ended at this step, else 0
        gamma:       discount factor

    Returns:
        (T,) array of TD errors
    """
    # YOUR CODE HERE
    pass
Loading docs…

The AI Mentor needs an account

It reads your code and the failing tests and nudges you toward the fix without handing you the answer. Free accounts get it on every problem you're working on today.