Epsilon-Greedy Action Selection
~15 mincode completion
Implement epsilon_greedy(q_values, epsilon, seed). Use for reproducibility. Return the chosen action index.
Examples
epsilon=0: always greedy (argmax)
- Input
- epsilon_greedy([1, 5, 2], 0, 0)
- Output
- 1
epsilon=1: always random (check within range)
- Input
- epsilon_greedy([0, 0, 0, 0], 1, 42)
- Output
- 2
epsilon=0: greedy regardless of seed
- Input
- epsilon_greedy([3, 1, 2], 0, 99)
- Output
- 0
Hints
Hint 1
You need the index of the extreme value, not the value itself.
Hint 2
Watch for this: always greedy regardless of epsilon.
Requirements
q_values: Q-value array, shape (n_actions,)epsilon: Exploration probability in [0, 1]: Random seed for reproducibility
Return Integer action index.
Constraints
Allowed library: NumPy only
Time limit: 200 ms, Memory: 64 MB
Where this shows up
~15 min
••••••
3 employers weight this skill
2 defense companies, 1 autonomy company. Top match scores 82.
Python
import numpy as np
def epsilon_greedy(q_values: np.ndarray, epsilon: float, seed: int) -> int:
"""
Select an action using the epsilon-greedy policy.
Args:
q_values: Q-value array, shape (n_actions,)
epsilon: Exploration probability in [0, 1]
seed: Random seed for reproducibility
Returns:
Integer action index.
"""
# YOUR CODE HERE
pass