A community opera-workshop bot is fluent on a writing benchmark but produces insulting cast notes. What should the evaluation plan add?
Select an answer to reveal the explanation.
Short Explanation
Think of a bot that aces a writing quiz and still writes insulting cast notes. Add toxicity or safety scoring. Fluency and ROUGE do not measure harm.
Full Explanation
Evaluate harmfulness, not only fluency or overlap. Insulting cast notes fail a safety dimension even when a writing benchmark looks fine. ROUGE-only or skipping safety leaves that harm unmeasured.