Generalised Advantage Estimation

~16 mincode completion

Implement gae(deltas, dones, gamma, lam) returning one advantage per step.

Examples

lambda = 0 is the one-step estimate: the advantages are the TD errors

Input
gae([1, 2, 3], [0, 0, 1], 0.99, 0)
Output
[1, 2, 3]

lambda = 1 with no discount sums every future error

Input
gae([1, 1, 1], [0, 0, 1], 1, 1)
Output
[3, 2, 1]

A late reward propagates backwards, halving each step

Input
gae([0, 0, 8], [0, 0, 1], 0.5, 1)
Output
[2, 4, 8]

Hints

Hint 1

Loop a fixed number of times and update the running value each pass.

Hint 2

A common slip here: accumulated forwards instead of backwards from the end.

Requirements

  • deltas: (T,) TD errors

  • dones: (T,) 1 if the episode ended at this step, else 0

  • gamma: discount factor

  • lam: GAE lambda in [0, 1]

  • Return (T,) advantages

Constraints

  • Allowed library: NumPy only

  • Time limit: 200 ms, Memory: 64 MB

Python
import numpy as np


def gae(deltas, dones, gamma, lam):
    """
    Generalised advantage estimation from precomputed TD errors.

    Args:
        deltas: (T,) TD errors
        dones:  (T,) 1 if the episode ended at this step, else 0
        gamma:  discount factor
        lam:    GAE lambda in [0, 1]

    Returns:
        (T,) advantages
    """
    # YOUR CODE HERE
    pass
Loading docs…

The AI Mentor needs an account

It reads your code and the failing tests and nudges you toward the fix without handing you the answer. Free accounts get it on every problem you're working on today.