A volunteer fire-hall roster tool uses GenAI to flag “anomalies” in last night’s run, but many flags are not real defects. Which quality metric is most directly suffering?
Select an answer to reveal the explanation.
Short Explanation
Lots of false alarms means weak precision—the flags aren’t hitting the real anomalies often enough. Sounding fluent doesn’t fix that.
Full Explanation
Precision concerns correctness of outputs for a specific objective, such as generated findings that truly identify anomalies. It is distinct from recall, fluency, load time, and token volume.