A four-node AHV cluster using RF2 loses one node to a failed drive. VMs continue to run, and the remaining nodes report normal CPU and network usage, but the cluster health score drops. What best explains the score change?
Select an answer to reveal the explanation.
Short Explanation
Think of the health score like a redundancy dashboard, not just a “VMs are on” light. When a node goes dark, you lose another place for data and cluster services to live, so the score drops even if workloads still run. Don't be fooled by running VMs; availability and redundancy can still be compromised.
Full Explanation
The cluster health score summarizes overall condition, not merely whether VMs are powered on. In a distributed AHV cluster, each node contributes CVM services, storage capacity, and fault-tolerance headroom. When one node becomes unavailable, the cluster may still serve VMs, but it has fewer active nodes, fewer replica placements, and less protection against a second failure. That reduction in availability and redundancy explains the lower score even when CPU, network, and VM power appear normal. A high CPU or memory load on the remaining nodes can create alerts or performance concerns, but the health score is not defined as a simple threshold that drops only after a fixed utilization percentage. Similarly, free storage capacity is one capacity dimension, but the score is not limited to container free space; node health, redundancy, and service availability also matter. Treating the decline as an NCC artifact because VMs are running is also wrong: a cluster can remain available while being less redundant, and health scoring should reflect that degraded state. Exam caveat: do not equate running VMs with full cluster health; assess redundancy and active node loss. Operational check: review health dashboard and NCC alerts for the unavailable node, confirm replica placement and capacity headroom, and validate second-node failure resilience.