A client called me in April, six days into a support ticket that wasn't going anywhere. Their VCF upgrade from 5.2 to 9.0.1 was stuck at the precheck stage, and the vendor they'd downloaded the bundles from wasn't returning calls fast enough to matter.

Took about four minutes to find the actual problem once I had console access. They'd pulled only the 9.0.1.0 bundles through the VCF Download Tool. The upgrade path from 5.2.0.0 requires the intermediate 5.2.1.0 bundles first — it's a documented skip-upgrade precheck, and it blocks scheduling until the missing bundles are present. Nothing wrong with the environment itself. Nobody had told them the bundle order mattered, and the tool doesn't exactly announce it either.

Six days, for a bundle sequencing problem that a decent retainer catches before it becomes a ticket at all.

Break-Fix Is the Floor, Not the Value

Every managed service arrangement will tell you they respond when something breaks. That's not a differentiator anymore — it's the minimum bar for calling yourself a support provider. If break-fix response is the headline feature of your retainer, you're paying for a phone number, and phone numbers are cheap.

The actual value of a retainer shows up in what never becomes a ticket in the first place. That's a much harder thing to sell, because it's invisible when it's working. Nobody notices the incident that didn't happen.

What Falls Through When Nobody's Watching Between Tickets

vSAN capacity is the clearest example. Capacity utilization climbs steadily, nobody's watching the trend line, and eventually a host failure triggers a resync that the cluster doesn't have the headroom to absorb gracefully. The failure that triggers it isn't the real problem — the real problem is that capacity crossed a threshold three months earlier and nobody flagged it.

NSX-T controller cluster health is a quieter version of the same failure mode. A controller can lose quorum without throwing an alert that anyone's actually watching for, and the first symptom shows up somewhere unrelated — east-west traffic behaving strangely, DFW rule pushes taking longer than they should, nothing that immediately points back to controller health.

Proactive monitoring isn't a nice-to-have layered on top of support. It's the difference between catching a threshold before it's a problem and explaining, after the fact, why nobody was watching.

Certificates are the third one, and honestly the most avoidable. SDDC Manager and NSX-T certificates get issued at bring-up, everyone moves on to the next project, and the expiration date just sits there. I've seen environments where nobody could say with confidence when their SDDC Manager cert expired without logging in and checking — and by the time it comes up, it's usually because something already threw an authentication error.

The Upgrade Path Doesn't Forgive Skipped Bundles

VCF lifecycle management follows a fixed order. VCF Operations gets upgraded first, then SDDC Manager, then NSX, then vCenter, then ESXi — and that sequence isn't a suggestion, it's enforced by the prechecks. Broadcom's own documentation on the skip-upgrade path is explicit about this: if you're jumping release levels, the intermediate upgrade bundles have to be present, or scheduling gets blocked outright.

The failure mode I see most often with clients who don't have proactive lifecycle support isn't a bad upgrade. It's a stalled one. Someone tries to move two or three release levels at once, downloads only the target bundles, and the environment sits at a precheck failure until someone figures out what's missing. That's downtime measured in days of confusion, not minutes of actual work.

Signs Your Retainer Isn't Doing Lifecycle Work
  • You find out about a new VCF release from a vendor email, not from your support provider.
  • Nobody has told you what release train you're currently on relative to what's current.
  • The last patch cycle happened because something broke, not because it was scheduled.
  • You've never seen the interoperability matrix checked before a proposed upgrade.

HCX Gets Forgotten More Than Anything Else on This List

HCX appliances drift out of version currency quietly because migrations aren't a daily activity for most environments — the mesh gets stood up, workloads move, and then HCX just sits there for months until the next migration wave. Service mesh health can degrade in that window without anyone noticing, and the first sign of trouble is usually a failed migration mid-execution, at which point you're troubleshooting an appliance nobody's checked on since it was deployed.

What a Monthly Health Report Should Actually Say

"All green, no issues" isn't a health report. It's a status update with no information in it.

A report worth reading has capacity numbers with a trend attached — not just current vSAN utilization, but where it's headed and when it crosses a threshold that matters. It names certificate expiration dates individually rather than burying them in a general compliance statement. It states patch currency against the actual current release train, with whatever specifically is blocking the next upgrade if something is. And it says something about HCX appliance versions and service mesh health even in a month with no migrations planned, because that's exactly the kind of month HCX drift goes unnoticed.

The Question Worth Asking About Your Current Arrangement

Does it leave you more confident in your environment than you were three months ago, or about the same?

Can you name your current vSAN capacity trend without logging in to check right now?

Do you know your SDDC Manager and NSX-T certificate expiration dates offhand?

Has anyone told you what release train you're on relative to current, unprompted?

If the honest answer to most of those is "I'd have to go check," the retainer is functioning as a phone number. That might be fine, depending on what you're paying for it. It's just worth knowing which one you actually have.