A municipal millpond office copies a pretraining-scale learning rate and a long epoch count into a LoRA run on a short instruction list. Validation dies by dinner. What should change?
Select an answer to reveal the explanation.
Short Explanation
A pretraining-scale learning rate and a long epoch count on a short LoRA list will kill validation by dinner. Treat adaptation knobs as gentler: smaller steps, fewer passes, watch the curve. NCCL traffic and a mid-run NIM hop do not repair a harsh schedule.
Full Explanation
Learning rate and epoch count for LLM adaptation are experimental knobs, and they are usually gentler than pretraining recipes. A short instruction list overfits quickly when those knobs are copied from a pretraining run. Smaller steps, fewer passes, and a watched validation curve are the associate response. Collectives and live serving do not repair a harsh adaptation schedule.