Discounted Return

~12 mincode completion

Implement discounted_return(rewards, gamma) that returns the total discounted return .

Examples

gamma=1: sum of all rewards

Input
discounted_return([1, 1, 1, 1], 1)
Output
4

gamma=0: only first reward counts

Input
discounted_return([5, 10, 20], 0)
Output
5

gamma=0.9: 1 + 0.9 + 0.81 = 2.71

Input
discounted_return([1, 1, 1], 0.9)
Output
2.71

Hints

Hint 1

Use a matrix product rather than nested loops, and check which operand transposes.

Hint 2

Do not forget to discount factor. That step is easy to skip.

Requirements

  • rewards: 1D array of rewards [r_0, r_1, ..., r_{T-1}]

  • gamma: Discount factor in [0, 1]

  • Return G_0 = sum(gamma^t * r_t for t in range(T))

Constraints

  • Allowed library: NumPy only

  • Time limit: 200 ms, Memory: 64 MB

Where this shows up

~12 min

3 employers weight this skill

2 defense companies, 1 autonomy company. Top match scores 82.

Python
import numpy as np

def discounted_return(rewards: np.ndarray, gamma: float) -> float:
    """
    Compute the discounted return from time 0.

    Args:
        rewards: 1D array of rewards [r_0, r_1, ..., r_{T-1}]
        gamma:   Discount factor in [0, 1]

    Returns:
        G_0 = sum(gamma^t * r_t for t in range(T))
    """
    # YOUR CODE HERE
    pass
Loading docs…

The AI Mentor needs an account

It reads your code and the failing tests and nudges you toward the fix without handing you the answer. Free accounts get it on every problem you're working on today.