The PPO Clipped Objective
Implement ppo_clip_objective(logp_new, logp_old, advantages, epsilon) returning the scalar mean objective.
Examples
An unchanged policy has ratio 1, so the objective is the mean advantage
- Input
- ppo_clip_objective([0, 0], [0, 0], [1, 3], 0.2)
- Output
- 2
A good action pushed twice as likely is capped at 1 + epsilon
- Input
- ppo_clip_objective([0.6931471805599453], [0], [1], 0.2)
- Output
- 1.2
The same ratio on a bad action is not capped at all
- Input
- ppo_clip_objective([0.6931471805599453], [0], [-1], 0.2)
- Output
- -2
Hints
Hint 1
bounds an array in one call.
Hint 2
A common slip here: took the max of the two terms instead of the min.
Requirements
logp_new: (T,) log pi_new(a_t | s_t)logp_old: (T,) log pi_old(a_t | s_t), from when the batch was collectedadvantages: (T,) advantage estimatesepsilon: clip range, e.g. 0.2Return scalar objective to maximise
Constraints
Allowed library: NumPy only
Time limit: 200 ms, Memory: 64 MB
Try similar problems(4)
Reinforcement Learning: Rewards, Senses and PPO · ~16 min
Reinforcement Learning: Rewards, Senses and PPO · ~20 min
Reinforcement Learning: Rewards, Senses and PPO · ~18 min
Reinforcement Learning: Rewards, Senses and PPO · ~14 min
import numpy as np
def ppo_clip_objective(logp_new, logp_old, advantages, epsilon):
"""
Mean PPO-clip surrogate objective over a batch.
Args:
logp_new: (T,) log pi_new(a_t | s_t)
logp_old: (T,) log pi_old(a_t | s_t), from when the batch was collected
advantages: (T,) advantage estimates
epsilon: clip range, e.g. 0.2
Returns:
scalar objective to maximise
"""
# YOUR CODE HERE
pass