Discounted Return
~12 mincode completion
Implement discounted_return(rewards, gamma) that returns the total discounted return G0.
Examples
gamma=1: sum of all rewards
- Input
- discounted_return([1, 1, 1, 1], 1)
- Output
- 4
gamma=0: only first reward counts
- Input
- discounted_return([5, 10, 20], 0)
- Output
- 5
gamma=0.9: 1 + 0.9 + 0.81 = 2.71
- Input
- discounted_return([1, 1, 1], 0.9)
- Output
- 2.71
Hints
Hint 1
Use a matrix product rather than nested loops, and check which operand transposes.
Hint 2
Do not forget to discount factor. That step is easy to skip.
Requirements
rewards: 1D array of rewards [r_0, r_1, ..., r_{T-1}]gamma: Discount factor in [0, 1]Return G_0 = sum(gamma^t * r_t for t in range(T))
Constraints
Allowed library: NumPy only
Time limit: 200 ms, Memory: 64 MB
Where this shows up
~12 min
••••••
3 employers weight this skill
2 defense companies, 1 autonomy company. Top match scores 82.
Python
import numpy as np
def discounted_return(rewards: np.ndarray, gamma: float) -> float:
"""
Compute the discounted return from time 0.
Args:
rewards: 1D array of rewards [r_0, r_1, ..., r_{T-1}]
gamma: Discount factor in [0, 1]
Returns:
G_0 = sum(gamma^t * r_t for t in range(T))
"""
# YOUR CODE HERE
pass