Repository navigation
kvm: dedicated live migration network with TLS, parallel streams and cluster CPU baseline - #14347
Open
nagaboinaramgopal wants to merge 1 commit into
Open
nagaboinaramgopal wants to merge 1 commit into
nagaboinaramgopal wants to merge 1 commit into
Conversation
…cluster CPU baseline Carry KVM live migration traffic on a dedicated network instead of the management network. A Migration traffic type is added to the zone's physical network with a per hypervisor label; each KVM host resolves its own IP on the labelled NIC and reports it, and the libvirt data stream is routed to it only when both the source and the destination host have one, otherwise migration falls back to the existing behaviour unchanged. listHosts returns each host's migration IP, shown on the host detail view, so an operator can confirm the dedicated network is in use. Migrations can be encrypted with libvirt native TLS and parallelised with multiple file-descriptor streams. Both are opt-in and gated on the host libvirt version, failing closed with a clear error when the host is too old rather than silently downgrading. Add a cluster-scoped CPU baseline model so a mixed cluster can present a uniform guest CPU, keeping a VM's CPU definition stable as it migrates across hosts. The baseline guest is pinned with fallback=forbid so it refuses a host that cannot provide the exact model rather than silently degrading the CPU the guest sees; the forbid is scoped to the baseline and leaves a VM with its own explicit CPU mode or model untouched. The baseline is validated against every host in the cluster before it is accepted, and hosts that cannot run the model are reported back instead of failing the migration later. Setting the baseline to "auto" computes the common-denominator model of the cluster's reachable hosts, by collecting each host's CPU and running cpu-baseline, and persists the computed model; it does not re-compute as hosts change, so the committed baseline stays stable. Unit tests cover the IP resolver (both-ends gating, non-KVM skip, blank values), the libvirt flag and typed-parameter construction, the version gate, the CPU baseline injection, the forbid fallback emission and scoping, the auto-compute and model parsing, and the host compatibility check including the fail-closed path on an unknown model. A Marvin smoke test exercises the migration traffic type API.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
KVM live migration runs over the management network and in cleartext. This moves it onto a dedicated network, lets it be encrypted and parallelised, and adds a cluster CPU baseline so VMs stay migratable across mixed CPU generations.
The dedicated network is a new Migration traffic type on the zone's physical network, configured once with a per hypervisor label, exactly like the Storage traffic type. Each KVM host resolves its own IP on the labelled NIC and reports it to the management server, and the libvirt data stream is pinned to it only when both the source and destination host have one, otherwise migration behaves exactly as it does today, so it can be enabled one host at a time. CloudStack does not allocate the IPs: the operator gives the migration NIC its address the same way the management NIC gets one, and CloudStack reads it from the labelled interface. listHosts returns each host's migration IP, shown on the host detail view, so an operator can confirm the dedicated network is in use. Removing the Migration traffic type clears the recorded IPs so migration stops targeting a decommissioned network.
Native TLS encryption and multiple file descriptor parallel streams sit on top, both opt-in. TLS sets the libvirt TLS flag; parallel streams set a configurable connection count for high latency links. Both are gated on the host libvirt version and fail closed with a clear error when the host is too old, rather than silently downgrading to an unencrypted or single stream migration.
cluster.cpu.baseline.model (cluster scoped) pins every user VM in the cluster to a chosen custom CPU model, so the guest sees the same CPU on every host and keeps migrating across mixed generations. The model is validated against every host in the cluster when it is set; a host that cannot run it is reported back instead of the migration failing later. The baseline guest is pinned with match='exact', so a host that cannot provide the model's features refuses the VM rather than the guest quietly receiving a different CPU, and with fallback='forbid' so libvirt will not substitute a near match model. This is scoped to the baseline; a VM with its own explicit CPU mode or model is left untouched. Setting the baseline to 'auto' computes the common denominator model of the cluster's reachable hosts and persists it; it does not recompute as hosts change, so the committed baseline stays stable. The baseline is set from a Configure CPU baseline action on the cluster, with a model picker that includes auto.
No existing behaviour changes unless an operator configures a migration network, enables TLS or parallel streams, or sets a cluster baseline. No database schema change is required.
Types of changes
Feature/Enhancement Scale or Bug Severity
Feature/Enhancement Scale
How Has This Been Tested?
Unit tests, all green:
They cover the Migration label delivery to the host and the agent resolving it to a local IP, the host detail being recorded on connect and cleared when the traffic type is removed, the both ends gate on the migration IP (non KVM skipped, blank rejected, resolved only when both hosts have one), the libvirt flag and typed parameter construction for TLS and parallel streams including the version gate, the CPU baseline injection at VM start (user instances only), the forbid fallback emission and its scoping, the auto compute and model parsing, and the host compatibility check including the fail closed path on an unknown model.
Verified end to end on a two host KVM cluster. A Migration traffic type was added to the physical network with the dedicated NIC as its label; each host resolved its own IP on that NIC and listHosts showed it. Migrating a running VM then moved the guest RAM over the migration interface (about 54 MB of pages captured on that NIC during the migration), while the management interface carried no migration traffic. The CPU baseline was set to a custom model and the running VM's definition was hard pinned (custom, match=exact, fallback=forbid) and migrated across both hosts; 'auto' computed the common model of the two hosts and stored it; an unsupported model was rejected at config time with the list of hosts that could not run it.
A Marvin smoke test (test/integration/smoke/test_migration_network.py) exercises adding the Migration traffic type and the cluster CPU baseline setting.
How did you try to break this feature and the system with this change?