Running Observation Normalisation
Implement welford_update(mean, m2, count, x) for one new scalar sample, returning the array [new_mean, new_m2, new_count].
Examples
The very first sample: the mean is that sample, spread is still zero
- Input
- welford_update(0, 0, 0, 5)
- Output
- [5, 0, 1]
A second sample of 7: mean 6, and M2 of 2 is a variance of 1
- Input
- welford_update(5, 0, 1, 7)
- Output
- [6, 2, 2]
A third sample of 3 pulls the mean back to 5
- Input
- welford_update(6, 2, 2, 3)
- Output
- [5, 8, 3]
Hints
Hint 1
Convert the input with before doing elementwise work.
Hint 2
A common slip here: updated m2 with the old mean twice instead of old then new.
Requirements
: running mean so far
m2: running sum of squared deviations so farcount: how many samples that summarisesx: the new sampleReturn (3,) array [new_mean, new_m2, new_count]
Constraints
Allowed library: NumPy only
Time limit: 200 ms, Memory: 64 MB
Try similar problems(4)
Reinforcement Learning: Rewards, Senses and PPO · ~12 min
Reinforcement Learning: Rewards, Senses and PPO · ~14 min
Reinforcement Learning: Rewards, Senses and PPO · ~14 min
Reinforcement Learning: Rewards, Senses and PPO · ~12 min
import numpy as np
def welford_update(mean, m2, count, x):
"""
One step of Welford's online mean and variance.
Args:
mean: running mean so far
m2: running sum of squared deviations so far
count: how many samples that summarises
x: the new sample
Returns:
(3,) array [new_mean, new_m2, new_count]
"""
# YOUR CODE HERE
pass