The PPO Clipped Objective

~18 mincode completion

Implement ppo_clip_objective(logp_new, logp_old, advantages, epsilon) returning the scalar mean objective.

Examples

An unchanged policy has ratio 1, so the objective is the mean advantage

Input
ppo_clip_objective([0, 0], [0, 0], [1, 3], 0.2)
Output
2

A good action pushed twice as likely is capped at 1 + epsilon

Input
ppo_clip_objective([0.6931471805599453], [0], [1], 0.2)
Output
1.2

The same ratio on a bad action is not capped at all

Input
ppo_clip_objective([0.6931471805599453], [0], [-1], 0.2)
Output
-2

Hints

Hint 1

bounds an array in one call.

Hint 2

A common slip here: took the max of the two terms instead of the min.

Requirements

  • logp_new: (T,) log pi_new(a_t | s_t)

  • logp_old: (T,) log pi_old(a_t | s_t), from when the batch was collected

  • advantages: (T,) advantage estimates

  • epsilon: clip range, e.g. 0.2

  • Return scalar objective to maximise

Constraints

  • Allowed library: NumPy only

  • Time limit: 200 ms, Memory: 64 MB

Python
import numpy as np


def ppo_clip_objective(logp_new, logp_old, advantages, epsilon):
    """
    Mean PPO-clip surrogate objective over a batch.

    Args:
        logp_new:   (T,) log pi_new(a_t | s_t)
        logp_old:   (T,) log pi_old(a_t | s_t), from when the batch was collected
        advantages: (T,) advantage estimates
        epsilon:    clip range, e.g. 0.2

    Returns:
        scalar objective to maximise
    """
    # YOUR CODE HERE
    pass
Loading docs…

The AI Mentor needs an account

It reads your code and the failing tests and nudges you toward the fix without handing you the answer. Free accounts get it on every problem you're working on today.