Quiz 1 Question 6 of 20

During training of a very deep network, gradients for the earliest layers become vanishingly small so those weights barely update. What problem is being described?

Select an answer to reveal the explanation.

Motivation