Generalised Advantage Estimation
Implement gae(deltas, dones, gamma, lam) returning one advantage per step.
Examples
lambda = 0 is the one-step estimate: the advantages are the TD errors
- Input
- gae([1, 2, 3], [0, 0, 1], 0.99, 0)
- Output
- [1, 2, 3]
lambda = 1 with no discount sums every future error
- Input
- gae([1, 1, 1], [0, 0, 1], 1, 1)
- Output
- [3, 2, 1]
A late reward propagates backwards, halving each step
- Input
- gae([0, 0, 8], [0, 0, 1], 0.5, 1)
- Output
- [2, 4, 8]
Hints
Hint 1
Loop a fixed number of times and update the running value each pass.
Hint 2
A common slip here: accumulated forwards instead of backwards from the end.
Requirements
deltas: (T,) TD errorsdones: (T,) 1 if the episode ended at this step, else 0gamma: discount factorlam: GAE lambda in [0, 1]Return (T,) advantages
Constraints
Allowed library: NumPy only
Time limit: 200 ms, Memory: 64 MB
Try similar problems(4)
Reinforcement Learning: Rewards, Senses and PPO · ~18 min
Reinforcement Learning: Rewards, Senses and PPO · ~20 min
Reinforcement Learning: Rewards, Senses and PPO · ~18 min
Reinforcement Learning: Rewards, Senses and PPO · ~14 min
import numpy as np
def gae(deltas, dones, gamma, lam):
"""
Generalised advantage estimation from precomputed TD errors.
Args:
deltas: (T,) TD errors
dones: (T,) 1 if the episode ended at this step, else 0
gamma: discount factor
lam: GAE lambda in [0, 1]
Returns:
(T,) advantages
"""
# YOUR CODE HERE
pass