Introduction
Day-2 in most infrastructure conversations gets reduced to one word: patching. VCF 9.1 treats it as two separate, ongoing jobs that happen to share a control plane. The first is applying fixes without breaking anything that’s running. The second is proving, continuously, that what’s running still matches what you intended to deploy. Broadcom built genuinely different tooling for each, and conflating them is a good way to think you’re covered on one when you’ve only handled the other.
Patching Isn’t One Workflow, It’s Three
VCF 9.1’s patching architecture starts from a fact that’s easy to forget mid-incident: not every component carries the same disruption risk if something goes wrong while patching it. Broadcom’s own engineering team splits the entire stack into three layers, each with a patching mechanism matched to what that layer can tolerate:
-> Management layer (VCF Management Services, VCF Operations, VCF Automation, Cloud Proxy, VCF Operations for Networks): architecturally separate from anything running workloads, so patching here carries no workload risk at all. This layer runs on a declarative model, you define a target version, and the Fleet Lifecycle service orchestrates the rest across the fleet, replacing what used to be manual, error-prone version-by-version steps.
-> Control plane layer (vCenter, NSX Manager, vSphere Supervisor, VKS): everything here has to stay available during the patch, since every other operation depends on it. The mechanism is matched to the patch type: vCenter Quick Patch handles security and minor fixes, Reduced Downtime Upgrade handles full version transitions, and vSphere Supervisor and VKS clusters get rolling updates. NSX keeps at least two manager nodes active throughout so management never actually drops.
-> Data plane layer (ESX, vSAN, NSX Edge): this is where workloads actually run, and it carries the strictest constraint of the three, patch the host without disrupting what’s on it. ESX Live Patch applies fixes directly in memory: no maintenance window, no evacuation, no reboot, and as of 9.1 this now extends to TPM-enabled hosts. When a reboot genuinely can’t be avoided, Quick Boot, pre-staging, and live vMotion evacuation keep the actual impact as small as possible.
Across all three layers, the same two disciplines apply: prechecks confirm a patch is likely to succeed before it commits to anything, and recoverable, migration-based designs give you a fallback when a precheck was wrong.
Declarative Lifecycle Management, in Practice
The mental shift underneath all three layers: you stop issuing a sequence of individual patch commands and start declaring a target version for the environment. VCF Operations, through the Fleet Lifecycle service, works out and executes the actual sequence needed to get there. This isn’t just a UX simplification, it’s what makes the scale numbers in 9.1 possible: a single instance now supports up to 5,000 ESX hosts, and parallel upgrade capacity has quadrupled to 256 clusters at once. Neither of those numbers is reachable if a human is still sequencing individual patch steps by hand.
Compliance Is a Separate, Continuous Job
Patched doesn’t mean compliant. A host can be running the latest build and still have drifted away from the configuration baseline you actually intended, a changed MOTD, a modified role permission, a setting an engineer tweaked during troubleshooting and never reverted. VCF Operations treats catching that drift as its own workflow, not a side effect of patching:
-> Configuration Management (the current name for what was previously called Configuration Drifts) lets you build configuration templates for vCenter and cluster objects, then schedule drift detection against them on a recurring basis rather than checking manually. For vSphere Configuration Profile-enabled clusters specifically, you get aggregated drift status across the whole fleet in one view.
-> Drift reports are downloadable as PDFs, scoped to whichever vCenter instances and retention period you choose, useful as an actual audit artifact rather than something you have to screenshot.
-> Git integration makes source control the template’s owner. Once you connect a Git repository, VCF Operations treats Git as authoritative for template versioning rather than the UI itself, which matters if you’re trying to run configuration-as-code discipline across a fleet rather than one-off manual edits.
-> VMware Advanced Cyber Compliance (ACC) goes a step further than detection. Built on VMware Salt, it continuously monitors your VCF 9.1 stack against a chosen security baseline, either a Security Configuration Guide (SCG) or PCI-DSS, and when an ESX host drifts out of compliance, ACC can automatically remediate it back to the desired state rather than just flagging it for someone to fix later.
Architectural Overview
What’s Actually New Architecturally in 9.1
Two structural changes are worth knowing if you’re coming from 9.0: the standalone VCF Operations Fleet Management Appliance no longer exists as its own component, its functionality (service registry, imported certificates, certificate signing requests) is absorbed into VCF Operations 9.1 directly, with the Fleet Lifecycle component taking over orchestration duties. And licensing moved out of VCF Operations entirely into a dedicated license server, installed automatically alongside VCF Operations 9.1 rather than living inside it.
This is the same pattern this series has already flagged once with the Identity Broker’s Appliance-to-Instance consolidation, VCF 9.1 is consistently folding what used to be standalone appliances into shared, unified services rather than just adding features to the existing ones.
Patch Sequencing Gets Stricter in 9.1.1
The layered model above still holds exactly as described, but VCF 9.1.1’s own release notes add a hard sequencing rule worth knowing before your next maintenance window: when moving from 9.1.0.x to 9.1.1, you patch the VCF Management Services Fleet Lifecycle component first, before touching anything else. VCF Operations itself cannot be patched in parallel with other components, and before patching ESX hosts you must first patch the VCF Operations instance and the license servers connected to it.
None of this changes the three-layer model, it reinforces it: Management layer goes first because everything downstream depends on Fleet Lifecycle being current. It does mean “patch in whatever order is convenient” is no longer a safe assumption at 9.1.1, the order is now enforced, not just recommended.
Choosing Where to Look First: A Quick Checklist
- Just deployed or upgraded to 9.1? Confirm your Fleet Lifecycle component is current before patching anything else, dependencies flow from it outward.
- Patching 9.1.0.x to 9.1.1? Patch order is enforced, not optional: VCF Management Services Fleet Lifecycle first, then VCF Operations and its connected license servers, before ESX hosts.
- Planning a patch cycle? Know which layer you’re touching before you start, the acceptable disruption and the mechanism both change by layer.
- Haven’t set up drift detection yet? It’s a separate configuration step from patching, being current on patches tells you nothing about configuration drift.
- Running a regulated environment? Advanced Cyber Compliance’s auto-remediation is worth deliberately choosing, or deliberately declining, rather than leaving on the default.
What’s Next
Next in this series: DR & Ransomware Recovery, Isolated Recovery, SRM, and VPC Isolation.
Further Reading (Official Broadcom Documentation)
- Faster Security Patching with Fewer Disruptions in VCF 9.1
- Scale, Simplify, and Secure Your Private Cloud Operations with VCF 9.1
- Securing your VMware Cloud Foundation 9.1 Environment
- Fleet Management: Configuration Management Overview
- VCF Operations 9.1.0.0 What’s New
- VMware Cloud Foundation 9.1.1.0 Release Notes, Getting to 9.1.1
