Visitar URL original
kvm: dedicated live migration network with TLS, parallel streams and cluster CPU baseline by nagaboinaramgopal · Pull Request #14347 · apache/cloudstack · GitHub
Skip to content

kvm: dedicated live migration network with TLS, parallel streams and cluster CPU baseline - #14347

Open
nagaboinaramgopal wants to merge 1 commit into
apache:mainfrom
nagaboinaramgopal:pr/live-migration-enhancements
Open

nagaboinaramgopal wants to merge 1 commit into
apache:mainfrom
nagaboinaramgopal:pr/live-migration-enhancements

Conversation

@nagaboinaramgopal

Copy link
Copy Markdown
Contributor

Description

KVM live migration runs over the management network and in cleartext. This moves it onto a dedicated network, lets it be encrypted and parallelised, and adds a cluster CPU baseline so VMs stay migratable across mixed CPU generations.

The dedicated network is a new Migration traffic type on the zone's physical network, configured once with a per hypervisor label, exactly like the Storage traffic type. Each KVM host resolves its own IP on the labelled NIC and reports it to the management server, and the libvirt data stream is pinned to it only when both the source and destination host have one, otherwise migration behaves exactly as it does today, so it can be enabled one host at a time. CloudStack does not allocate the IPs: the operator gives the migration NIC its address the same way the management NIC gets one, and CloudStack reads it from the labelled interface. listHosts returns each host's migration IP, shown on the host detail view, so an operator can confirm the dedicated network is in use. Removing the Migration traffic type clears the recorded IPs so migration stops targeting a decommissioned network.

Native TLS encryption and multiple file descriptor parallel streams sit on top, both opt-in. TLS sets the libvirt TLS flag; parallel streams set a configurable connection count for high latency links. Both are gated on the host libvirt version and fail closed with a clear error when the host is too old, rather than silently downgrading to an unencrypted or single stream migration.

cluster.cpu.baseline.model (cluster scoped) pins every user VM in the cluster to a chosen custom CPU model, so the guest sees the same CPU on every host and keeps migrating across mixed generations. The model is validated against every host in the cluster when it is set; a host that cannot run it is reported back instead of the migration failing later. The baseline guest is pinned with match='exact', so a host that cannot provide the model's features refuses the VM rather than the guest quietly receiving a different CPU, and with fallback='forbid' so libvirt will not substitute a near match model. This is scoped to the baseline; a VM with its own explicit CPU mode or model is left untouched. Setting the baseline to 'auto' computes the common denominator model of the cluster's reachable hosts and persists it; it does not recompute as hosts change, so the committed baseline stays stable. The baseline is set from a Configure CPU baseline action on the cluster, with a model picker that includes auto.

No existing behaviour changes unless an operator configures a migration network, enables TLS or parallel streams, or sets a cluster baseline. No database schema change is required.

Types of changes

  • New feature (non-breaking change which adds functionality)

Feature/Enhancement Scale or Bug Severity

Feature/Enhancement Scale

  • Major

How Has This Been Tested?

Unit tests, all green:

NetworkServiceImplTest                          Tests run: 73
NetworkOrchestratorTest                         Tests run: 66
LibvirtComputingResourceTest                    Tests run: 324
VirtualMachineManagerImplMigrationIpTest        Tests run: 8
VirtualMachineManagerImplCpuCompatTest          Tests run: 6
VirtualMachineManagerImplCpuBaselineTest        Tests run: 5
MigrateKVMAsyncTest                             Tests run: 15
LibvirtMigrateCommandWrapperTest                Tests run: 39
LibvirtCheckNetworkCommandWrapperTest           Tests run: 2
LibvirtCheckCpuCompatibilityCommandWrapperTest  Tests run: 10
LibvirtBaselineCpuCommandWrapperTest            Tests run: 3
LibvirtGetHostCpuModelCommandWrapperTest        Tests run: 2

They cover the Migration label delivery to the host and the agent resolving it to a local IP, the host detail being recorded on connect and cleared when the traffic type is removed, the both ends gate on the migration IP (non KVM skipped, blank rejected, resolved only when both hosts have one), the libvirt flag and typed parameter construction for TLS and parallel streams including the version gate, the CPU baseline injection at VM start (user instances only), the forbid fallback emission and its scoping, the auto compute and model parsing, and the host compatibility check including the fail closed path on an unknown model.

Verified end to end on a two host KVM cluster. A Migration traffic type was added to the physical network with the dedicated NIC as its label; each host resolved its own IP on that NIC and listHosts showed it. Migrating a running VM then moved the guest RAM over the migration interface (about 54 MB of pages captured on that NIC during the migration), while the management interface carried no migration traffic. The CPU baseline was set to a custom model and the running VM's definition was hard pinned (custom, match=exact, fallback=forbid) and migrated across both hosts; 'auto' computed the common model of the two hosts and stored it; an unsupported model was rejected at config time with the list of hosts that could not run it.

A Marvin smoke test (test/integration/smoke/test_migration_network.py) exercises adding the Migration traffic type and the cluster CPU baseline setting.

How did you try to break this feature and the system with this change?

  • Migration IP on only one of the two hosts: migration falls back to the management path (both ends gate), confirmed by unit test and on the cluster.
  • TLS and parallel streams enabled on a host whose libvirt is too old: migration fails closed with a clear error instead of downgrading.
  • A bogus and a well formed but unsupported baseline model: config is rejected, the value is not stored, and the unknown model path fails closed rather than passing an unchecked string to the hypervisor.
  • Removing the Migration traffic type after it was in use: the recorded host migration IPs are cleared so later migrations do not target the decommissioned network.
  • A transient failure recording the migration IP during host connect: it is logged and the host still connects, rather than the informational detail failing the connection.
migration-traffic-type-dropdown cpu-baseline-dialog migration-dropdown-cropped

…cluster CPU baseline

Carry KVM live migration traffic on a dedicated network instead of the
management network. A Migration traffic type is added to the zone's physical
network with a per hypervisor label; each KVM host resolves its own IP on the
labelled NIC and reports it, and the libvirt data stream is routed to it only
when both the source and the destination host have one, otherwise migration
falls back to the existing behaviour unchanged. listHosts returns each host's
migration IP, shown on the host detail view, so an operator can confirm the
dedicated network is in use.

Migrations can be encrypted with libvirt native TLS and parallelised with
multiple file-descriptor streams. Both are opt-in and gated on the host
libvirt version, failing closed with a clear error when the host is too old
rather than silently downgrading.

Add a cluster-scoped CPU baseline model so a mixed cluster can present a
uniform guest CPU, keeping a VM's CPU definition stable as it migrates across
hosts. The baseline guest is pinned with fallback=forbid so it refuses a host
that cannot provide the exact model rather than silently degrading the CPU the
guest sees; the forbid is scoped to the baseline and leaves a VM with its own
explicit CPU mode or model untouched. The baseline is validated against every
host in the cluster before it is accepted, and hosts that cannot run the model
are reported back instead of failing the migration later.

Setting the baseline to "auto" computes the common-denominator model of the
cluster's reachable hosts, by collecting each host's CPU and running
cpu-baseline, and persists the computed model; it does not re-compute as hosts
change, so the committed baseline stays stable.

Unit tests cover the IP resolver (both-ends gating, non-KVM skip, blank
values), the libvirt flag and typed-parameter construction, the version gate,
the CPU baseline injection, the forbid fallback emission and scoping, the
auto-compute and model parsing, and the host compatibility check including the
fail-closed path on an unknown model. A Marvin smoke test exercises the
migration traffic type API.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant