Quiz 9 Question 2 of 20

A CIO at a multi-strategy hedge fund is selecting between three foundation models to power an earnings call summarization and sentiment classification system. The vendor benchmarks show strong scores on MMLU, HellaSwag, and HumanEval. Before deploying to production, the risk team requires rigorous domain-specific evaluation. Which evaluation methodology most accurately predicts real-world performance for this specific financial NLP use case?

Select an answer to reveal the explanation.

Motivation