advanced 3 min answer

Eleven weeks before a licence audit response is due, you learn that the database vendor's metric counts every physical core in the cluster the VM is permitted to run on, not the 8 vCPUs you assigned it. The cluster is three hosts of 64 cores. Sequence the remediation under live traffic and name the point of no return.

licensingconstraintsauditvirtualisationcutover
Show the full answer Hide the answer

The number before anything else

Three hosts at 64 cores is 192 physical cores inside the licensable boundary against the 8 vCPUs the design assumed, before any core factor is applied. Get the exact figure out of the contract's own metric definition and core factor table before designing anything, because the remediation is chosen by which number it removes, and the gap is usually one to two orders of magnitude rather than a percentage.

Two facts drive every step below. Licence metrics of this shape count where a VM may run, not where it is running, so live-migration scope is the architectural variable. And most of them count software that is installed, whether or not it is serving anything in production.

The sequence

  1. Read the vendor's own definition of the partitioning it recognises. The remediation has to match that text. A general argument about CPU affinity or pinning is not a defence if the policy names specific mechanisms.
  2. Build a licensable boundary you are willing to pay for: a dedicated host or a cluster configured so the database VM cannot migrate outside it. Capture the configuration as dated evidence at the moment you make it, because the audit question is about the period, not about today.
  3. Replicate into the new boundary and move read traffic first, behind a flag, with the old primary still writable. This validates the host, the storage path and the backup job under real load without a cutover.
  4. Cut writes over in a window, and keep reverse replication running for two weeks. Failback has to be a configuration change, not a restore.
  5. Remove the engine binaries from every host still in the old cluster, and re-export the install inventory.
  6. Assemble the audit pack: cluster policy export, install inventory, and the dates each became true.

Where data can diverge, and how you would know

During step 3 the divergence is replication lag: any report run against the replica inside the lag window disagrees with the primary, and finance reports are the ones that get noticed. Instrument lag, alert on it above the window your reports tolerate, and label replica-sourced reports. During step 4 the risk is split writes from a client that caches the old endpoint, which connection-level audit logging on the old primary will show and application logs will not.

The point of no return

Uninstalling the binaries from the old hosts. Before that, every step is reversible with a flag or a failback. After it, reinstalling to fail back re-creates a countable installation and destroys the dated inventory the audit response rests on. Schedule the uninstall after the two-week failback window has expired, not before, and make it a separate approved change.

When this is the wrong answer

Price the alternative first. If the exposure is small enough, buying the licences is cheaper and far faster than eleven weeks of topology work by people who have other commitments, and the audit response is a cheque rather than an architecture. If the estate is already heading to a managed service within a year, a negotiated interim position usually beats a rushed change that has to be unwound. The architecture move is justified when the annual recurring exposure exceeds the remediation cost, which is a calculation and not a judgement call.