Allowing capacity for host maintenance and failures
Cluster capacity should include room for a host to be unavailable. If every host is fully allocated, a maintenance event or failure may leave too little capacity for the remaining workload. The required reserve depends on the availability design and application priorities.
Plan beyond the nominal total
- Define how many host failures or maintenance events the design must tolerate and which workloads take priority.
- Size remaining CPU and RAM for that scenario, including hypervisor overhead and peak demand.
- Check storage, network and licence constraints as well as compute capacity.
- Agree workload placement, maintenance sequencing and acceptance tests for the intended recovery behaviour.
Do not calculate usable production capacity by simply adding the labels on all hosts. Ask for both total resources and the capacity available under the agreed failure scenario. Review the reserve again when adding virtual machines or increasing allocations.