Skip to content

Adopt Detached VMware VMDKs as FCD Volumes - #356

Draft
anokfireball wants to merge 3 commits into
stable/2025.1-m3from
manage-ephemeral-to-fcd-2025
Draft

Adopt Detached VMware VMDKs as FCD Volumes#356
anokfireball wants to merge 3 commits into
stable/2025.1-m3from
manage-ephemeral-to-fcd-2025

Conversation

@anokfireball

@anokfireball anokfireball commented Aug 3, 2026

Copy link
Copy Markdown
Member

This is the stable/2025.1-m3 version of #349.

Problem

As part of our VMware offramp, Nova will have to move some image-backed servers from VMware to KVM. These servers boot from Nova-owned VMDKs on VMware ephemeral datastores. KVM cannot use those VMDKs directly. Nova first needs a Cinder volume that it can attach during the move.

Solution

This PR does not create the final NetApp-driver-owned volume that might be expected for KVM. It instead helps create an intermediate FCD volume that can be attached to both VMware and KVM via #288. A successive user-initiated volume retype will move that volume to the final backend.

The VMware FCD driver gets a manage_existing path for detached VMDKs (also transparently supporting existing FCDs on ephemeral datastores). The caller passes a reference like:

{
  "source-name": "[datastore] path/to/disk.vmdk",
  "size_gb": 64
}

If Nova already knows the vSphere FCD id, it can also pass source-id. Cinder will then skip the registration of the FCD and reuse the existing one as-is.

The driver then:

  1. resolves the source datastore by name through vCenter
  2. registers the detached VMDK as an FCD, or reuses the FCD from source-id
  3. moves the FCD to a normal Cinder FCD datastore if needed
  4. returns the final FCD provider_location

The source datastore is only used as a one-time source location through vCenter. Cinder does not treat it as managed capacity.

Recovery behavior

The driver does not delete or unregister the disk during adoption. If adoption fails before relocation, the driver marks the volume as error. The resize fails, but Nova can cleanly recover the source VM by reattaching the original VMDK to the VMware server. If adoption fails during or after relocation, the volume stays in error_managing. At that point operator intervention is required.

Policy

volume_extension:volume_manage now accepts service-token callers through is_service_request:True. This lets Nova call manage_existing on behalf of the user without requiring the user to have admin permisssion. Deployments with a policy.yaml override for volume_extension:volume_manage must add or is_service_request:True.

Let drivers signal safe pre-relocate failures by setting volume
status to 'error' so the manage_existing flow preserves that state
on rollback instead of always reverting to 'error_managing'. Add
recovery logging in ManageExistingTask when the DB update cannot
be persisted after a driver has already adopted the backend object.

Change-Id: I4d060160561fd8df4e1869194a14934edf40f060
Add manage_existing and manage_existing_get_size overrides to
VMwareVStorageObjectDriver. Parses the '[ds] path.vmdk' source-name
ref, validates caller-provided size_gb, registers the VMDK as a
First Class Disk via RegisterDisk, and relocates it to the target
datastore when source and target moref differ.

All pre-relocate failures mark the volume as 'error' (safe for Nova
abort); relocate and post-relocate failures leave status at the
flow-owned 'error_managing'. No error path calls delete_fcd or
unregister_disk.

Change-Id: Ic7faba0b8aefaff1472a5ff0d380547c856f1a63
Problem:
The volume_extension:volume_manage policy defaults to rule:admin_api,
so only admin users can call the manage_existing API.  When another
service (e.g. Nova) calls manage_existing on behalf of a user, the
request carries the user's token as primary (for correct volume
ownership and quota attribution) plus the service's token as
secondary.  The non-admin user token is rejected by the policy.

Solution:
Before enforcing the admin-only policy, check whether the request
carries a valid service token using the existing is_service_request()
helper (checks ctxt.service_roles against the configured
service_token_roles).  If the request comes from a trusted service,
skip the policy check.  This mirrors the pattern already used in
attachment_deletion_allowed() for service-initiated detach operations.

The admin path and the non-admin-without-service-token path remain
unchanged.

Change-Id: I88b07712f399dae7654d238605f6f410a1b5a9ba
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant