Quiz 14 Question 19 of 20

An engineer describes an RL agent that learns action-value estimates from experience gathered under an exploratory behavior policy, while the estimates themselves converge toward the optimal greedy policy. Which classic algorithm best matches this description?

Select an answer to reveal the explanation.

Motivation