Policy Entropy and the Cost of Exploring

~12 mincode completion

Implement gaussian_entropy(log_std) for a diagonal Gaussian policy, returning the scalar total entropy.

Examples

Three axes at sigma = 1: three times the constant

Input
gaussian_entropy([0, 0, 0])
Output
4.25682

One dimension at sigma = 1 is the constant itself

Input
gaussian_entropy([0])
Output
1.41894

A tighter policy explores less, and the entropy falls by 1 per axis

Input
gaussian_entropy([-1, -1, -1])
Output
1.25682

Hints

Hint 1

is the natural log, which is what this formula wants.

Hint 2

Watch for this: used the discrete shannon sum p log p on a continuous policy.

Requirements

  • log_std: (d,) log standard deviation per action dimension

  • Return scalar total entropy, summed over dimensions

Constraints

  • Allowed library: NumPy only

  • Time limit: 200 ms, Memory: 64 MB

Python
import numpy as np


def gaussian_entropy(log_std):
    """
    Differential entropy of a diagonal Gaussian policy.

    Args:
        log_std: (d,) log standard deviation per action dimension

    Returns:
        scalar total entropy, summed over dimensions
    """
    # YOUR CODE HERE
    pass
Loading docs…

The AI Mentor needs an account

It reads your code and the failing tests and nudges you toward the fix without handing you the answer. Free accounts get it on every problem you're working on today.