Returns to Go
Implement returns_to_go(rewards, gamma) returning one return per step, as a 1D array the same length as rewards.
Examples
No discount: every step still has the whole +10 ahead of it
- Input
- returns_to_go([0, 0, 10], 1)
- Output
- [10, 10, 10]
Halving the future: the goal is worth less the further back you stand
- Input
- returns_to_go([0, 0, 10], 0.5)
- Output
- [2.5, 5, 10]
A reward every step, gamma 0.9
- Input
- returns_to_go([1, 1, 1, 1], 0.9)
- Output
- [3.439, 2.71, 1.9, 1]
Hints
Hint 1
Loop a fixed number of times and update the running value each pass.
Hint 2
Watch for this: accumulated forwards so early steps got the smallest return.
Requirements
rewards: (T,) rewards, in time ordergamma: discount factor in [0, 1]Return (T,) array where entry t is r_t + gamma * r_{t+1} + ...
Constraints
Allowed library: NumPy only
Time limit: 200 ms, Memory: 64 MB
Try similar problems(4)
Reinforcement Learning: Rewards, Senses and PPO · ~14 min
Reinforcement Learning: Rewards, Senses and PPO · ~12 min
Reinforcement Learning: Rewards, Senses and PPO · ~14 min
Reinforcement Learning: Rewards, Senses and PPO · ~14 min
import numpy as np
def returns_to_go(rewards, gamma):
"""
Discounted return remaining at each step of one episode.
Args:
rewards: (T,) rewards, in time order
gamma: discount factor in [0, 1]
Returns:
(T,) array where entry t is r_t + gamma * r_{t+1} + ...
"""
# YOUR CODE HERE
pass