Policy Entropy and the Cost of Exploring
Implement gaussian_entropy(log_std) for a diagonal Gaussian policy, returning the scalar total entropy.
Examples
Three axes at sigma = 1: three times the constant
- Input
- gaussian_entropy([0, 0, 0])
- Output
- 4.25682
One dimension at sigma = 1 is the constant itself
- Input
- gaussian_entropy([0])
- Output
- 1.41894
A tighter policy explores less, and the entropy falls by 1 per axis
- Input
- gaussian_entropy([-1, -1, -1])
- Output
- 1.25682
Hints
Hint 1
is the natural log, which is what this formula wants.
Hint 2
Watch for this: used the discrete shannon sum p log p on a continuous policy.
Requirements
log_std: (d,) log standard deviation per action dimensionReturn scalar total entropy, summed over dimensions
Constraints
Allowed library: NumPy only
Time limit: 200 ms, Memory: 64 MB
Try similar problems(4)
Reinforcement Learning: Rewards, Senses and PPO · ~14 min
Reinforcement Learning: Rewards, Senses and PPO · ~12 min
Reinforcement Learning: Rewards, Senses and PPO · ~14 min
Reinforcement Learning: Rewards, Senses and PPO · ~14 min
import numpy as np
def gaussian_entropy(log_std):
"""
Differential entropy of a diagonal Gaussian policy.
Args:
log_std: (d,) log standard deviation per action dimension
Returns:
scalar total entropy, summed over dimensions
"""
# YOUR CODE HERE
pass