A $40B long-only equity manager wants to assess where it stands on the AI adoption curve relative to quantitative competitors. The CIO commissions an internal maturity review and discovers the firm has deployed ML-based factor screening but lacks systematic model governance, has no centralized feature store, and still uses Excel for a majority of risk aggregation. Which AI maturity framework dimension most directly identifies the gap between ad-hoc ML deployment and repeatable, governed AI production?
Select an answer to reveal the explanation.
Short Explanation and Infographic
Think of AI maturity like a pilot's checklist — having a plane and even knowing how to fly it isn't the same as having repeatable, documented pre-flight procedures that work the same way every time. The firm has the plane (ML models) but lacks the checklist infrastructure (MLOps governance). MLOps maturity — covering model versioning, monitoring, feature stores, and deployment pipelines — is the dimension that separates 'we built something that works' from 'we operate AI reliably at scale.' MLOps capability maturity is the correct lens here.
Full explanation below image
Full Explanation
AI maturity frameworks used in financial services — including variants of Google's MLOps Maturity Model, Gartner's AI Maturity Model, and elements codified in SR 11-7 for model risk — consistently identify MLOps capability as the hinge between experimentation and production-grade AI. The firm in this scenario exhibits classic Level 1 (ad-hoc) characteristics: models exist and produce outputs, but there is no centralized feature store (meaning features may be computed inconsistently across models), no model lifecycle governance (meaning champion/challenger regimes, version rollbacks, and drift monitoring are absent), and legacy tools remain in the critical risk path.
MLOps maturity directly addresses all three gaps. A feature store enforces consistent, auditable feature computation. Model governance frameworks — aligned with SR 11-7 model risk management — mandate documentation, validation, and ongoing monitoring. Replacing Excel in risk aggregation with governed pipelines eliminates fragility and supports regulatory examination. These are structural, repeatable capabilities, not one-off improvements.
Option A (data volume benchmarking) confuses data scale with operational maturity. A firm can have petabytes of data and still run models in ad-hoc notebooks. Option C (headcount ratio) is a talent metric, not an infrastructure or process metric — a small, well-governed team outperforms a large ungoverned one. Option D (number of live signals) measures output quantity, not the quality or repeatability of the processes that produce and sustain those signals. The NIST AI RMF's 'Manage' function similarly emphasizes that operational processes and governance structures — not model count — determine whether AI deployment is trustworthy and sustainable.