Quiz 13 Question 15 of 20

What is the primary objective of the policy-gradient method in reinforcement learning?

Select an answer to reveal the explanation.

Motivation