When Dying Early Pays
Implement compare_exits(step_cost, death_penalty, horizon, gamma) returning the array [survive_value, die_now_value].
Examples
Attempt 2's actual arithmetic: 700 steps of misery against one -10
- Input
- compare_exits(0.2, 10, 700, 1)
- Output
- [-150, -10]
No step cost, heavy discount: the same death is cheaper later
- Input
- compare_exits(0, 10, 2, 0.5)
- Output
- [-2.5, -10]
Step cost only: the discounted sum 1 + 0.5 + 0.25
- Input
- compare_exits(1, 0, 3, 0.5)
- Output
- [-1.75, 0]
Hints
Hint 1
Convert the input with before doing elementwise work.
Hint 2
Do not forget to discount the terminal penalty by gamma to the horizon. That step is easy to skip.
Requirements
step_cost: c, the penalty paid at every surviving step (>= 0)death_penalty: d, the penalty paid once, on death (>= 0)horizon: H, steps until the episode ends anywaygamma: discount factor in [0, 1]Return (2,) array [survive_value, die_now_value]
Constraints
Allowed library: NumPy only
Time limit: 200 ms, Memory: 64 MB
Try similar problems(4)
Reinforcement Learning: Rewards, Senses and PPO · ~14 min
Reinforcement Learning: Rewards, Senses and PPO · ~12 min
Reinforcement Learning: Rewards, Senses and PPO · ~14 min
Reinforcement Learning: Rewards, Senses and PPO · ~12 min
import numpy as np
def compare_exits(step_cost, death_penalty, horizon, gamma):
"""
Discounted value of surviving to the horizon against dying immediately.
Args:
step_cost: c, the penalty paid at every surviving step (>= 0)
death_penalty: d, the penalty paid once, on death (>= 0)
horizon: H, steps until the episode ends anyway
gamma: discount factor in [0, 1]
Returns:
(2,) array [survive_value, die_now_value]
"""
# YOUR CODE HERE
pass