A hatchery-incubator network's training loss oscillates and never settles, while instance CPU is idle and S3 throughput is fine. Someone wants a larger training instance. What is the real problem?
Select an answer to reveal the explanation.
Short Explanation
The hatchery loss oscillates and never settles, but CPU is idle and S3 is fine. That is a convergence issue, not a bigger instance. Translate is not a settling loss, and an oscillating curve is not a minimum.
Full Explanation
Oscillating or exploding loss with no stable minimum is a convergence failure. Idle CPU and healthy S3 throughput rule out ingest capacity, so a larger instance is the wrong first fix. Translate is not a convergence diagnosis.