Epsilon-Greedy Action Selection

~15 mincode completion

Implement epsilon_greedy(q_values, epsilon, seed). Use for reproducibility. Return the chosen action index.

Examples

epsilon=0: always greedy (argmax)

Input
epsilon_greedy([1, 5, 2], 0, 0)
Output
1

epsilon=1: always random (check within range)

Input
epsilon_greedy([0, 0, 0, 0], 1, 42)
Output
2

epsilon=0: greedy regardless of seed

Input
epsilon_greedy([3, 1, 2], 0, 99)
Output
0

Hints

Hint 1

You need the index of the extreme value, not the value itself.

Hint 2

Watch for this: always greedy regardless of epsilon.

Requirements

  • q_values: Q-value array, shape (n_actions,)

  • epsilon: Exploration probability in [0, 1]

  • : Random seed for reproducibility

  • Return Integer action index.

Constraints

  • Allowed library: NumPy only

  • Time limit: 200 ms, Memory: 64 MB

Where this shows up

~15 min

3 employers weight this skill

2 defense companies, 1 autonomy company. Top match scores 82.

Python
import numpy as np

def epsilon_greedy(q_values: np.ndarray, epsilon: float, seed: int) -> int:
    """
    Select an action using the epsilon-greedy policy.

    Args:
        q_values: Q-value array, shape (n_actions,)
        epsilon:  Exploration probability in [0, 1]
        seed:     Random seed for reproducibility

    Returns:
        Integer action index.
    """
    # YOUR CODE HERE
    pass
Loading docs…

The AI Mentor needs an account

It reads your code and the failing tests and nudges you toward the fix without handing you the answer. Free accounts get it on every problem you're working on today.