Potential-Based Reward Shaping
Implement shape_rewards(rewards, potentials, gamma), where potentials has one more entry than rewards (the potential of every state visited, including the last). Return the shaped rewards.
Examples
Undiscounted, a potential that climbs by 1 pays 1 a step
- Input
- shape_rewards([0, 0], [0, 1, 2], 1)
- Output
- [1, 1]
The same potentials at gamma 0.5 pay differently
- Input
- shape_rewards([0, 0], [0, 1, 2], 0.5)
- Output
- [0.5, 0]
A constant potential adds nothing at all, whatever the rewards
- Input
- shape_rewards([1, 2, 3], [5, 5, 5, 5], 1)
- Output
- [1, 2, 3]
Hints
Hint 1
Convert the input with before doing elementwise work.
Hint 2
Watch for this: used phi s minus phi s next the wrong way round.
Requirements
rewards: (T,) environment rewardspotentials: (T+1,) potential of each visited stategamma: discount factorReturn (T,) shaped rewards r_t + gamma * phi_{t+1} - phi_t
Constraints
Allowed library: NumPy only
Time limit: 200 ms, Memory: 64 MB
Try similar problems(4)
Reinforcement Learning: Rewards, Senses and PPO · ~14 min
Reinforcement Learning: Rewards, Senses and PPO · ~12 min
Reinforcement Learning: Rewards, Senses and PPO · ~14 min
Reinforcement Learning: Rewards, Senses and PPO · ~14 min
import numpy as np
def shape_rewards(rewards, potentials, gamma):
"""
Potential-based shaping, which leaves the optimal policy unchanged.
Args:
rewards: (T,) environment rewards
potentials: (T+1,) potential of each visited state
gamma: discount factor
Returns:
(T,) shaped rewards r_t + gamma * phi_{t+1} - phi_t
"""
# YOUR CODE HERE
pass