A city-council clerk wants replies that stay helpful and on-register. They can supervised-fine-tune on gold replies, run a human-feedback alignment experiment, or use SteerLM-style attribute control at inference. How should they choose?
Select an answer to reveal the explanation.
Short Explanation
Helpful, on-register replies. Match the recipe to the data they have: SFT when gold replies exist, a human-feedback run when preference ranks exist, SteerLM when they need inference-time attribute control. Alignment is not only a reward-model definition, and it is not a cooling study.
Full Explanation
SFT, a human-feedback or RLHF-style recipe, and SteerLM are alternative alignment experiments, not a single required path. The match depends on whether the office has gold replies, preference ranks, or a need for inference-time attribute control. That choice is an experiment-design item, not a later-domain safety-policy rewrite and not a datacenter study. Associate practice picks the recipe the data can actually support.