A glasshouse orchid desk renders the identical weekly care notice with two custom Riva TTS voices and plays both for interns. What success measure fits that TTS comparison?
Select an answer to reveal the explanation.
Short Explanation
Judging a spoken care notice with a picture-quality score is like grading a radio show with a paint chart. TTS customizations need speech-appropriate outcomes—listener preference or whether hearers catch the message. Hold the script fixed so the voice is what you compare.
Full Explanation
A fair TTS experiment holds the script constant and varies only the voice. Success should reflect speech quality for listeners—preference ratings or comprehension—not image metrics, cluster scale knobs, or GPU utilization. Matching the metric to the modality under test keeps the comparison valid.