Agents learn by interacting with an environment. Rewards shape behavior through the Markov Decision Process framework, the basis of game-playing AI.
Discounted Return
~12 min· Hard
Epsilon-Greedy Action Selection
~15 min· Hard
Q-Learning Update
~25 min· Medium
Sign in for the concept check
Optional multiple-choice questions on the ideas behind this section. Most useful after you have tried the coding problems above.