Potential-Based Reward Shaping

~12 mincode completion

Implement shape_rewards(rewards, potentials, gamma), where potentials has one more entry than rewards (the potential of every state visited, including the last). Return the shaped rewards.

Examples

Undiscounted, a potential that climbs by 1 pays 1 a step

Input
shape_rewards([0, 0], [0, 1, 2], 1)
Output
[1, 1]

The same potentials at gamma 0.5 pay differently

Input
shape_rewards([0, 0], [0, 1, 2], 0.5)
Output
[0.5, 0]

A constant potential adds nothing at all, whatever the rewards

Input
shape_rewards([1, 2, 3], [5, 5, 5, 5], 1)
Output
[1, 2, 3]

Hints

Hint 1

Convert the input with before doing elementwise work.

Hint 2

Watch for this: used phi s minus phi s next the wrong way round.

Requirements

  • rewards: (T,) environment rewards

  • potentials: (T+1,) potential of each visited state

  • gamma: discount factor

  • Return (T,) shaped rewards r_t + gamma * phi_{t+1} - phi_t

Constraints

  • Allowed library: NumPy only

  • Time limit: 200 ms, Memory: 64 MB

Python
import numpy as np


def shape_rewards(rewards, potentials, gamma):
    """
    Potential-based shaping, which leaves the optimal policy unchanged.

    Args:
        rewards:    (T,) environment rewards
        potentials: (T+1,) potential of each visited state
        gamma:      discount factor

    Returns:
        (T,) shaped rewards r_t + gamma * phi_{t+1} - phi_t
    """
    # YOUR CODE HERE
    pass
Loading docs…

The AI Mentor needs an account

It reads your code and the failing tests and nudges you toward the fix without handing you the answer. Free accounts get it on every problem you're working on today.