Versions Compared

Key

  • This line was added.
  • This line was removed.
  • Formatting was changed.

...

ScenariosCurrent behaviorHow to deal with it
  • Scenario 1: Connection failures between the agent and the management server
Agent cannot send Answer to management server, therefore timed out on management server

Agent: (for reconcile command only)

  1. Save command info as JSON file
  2. Update command state in JSON when start/process/complete the command
    1. STARTED
    2. PROCESSING
    3. COMPLETE/FAILED
  3. Save Answer in the JSON file
  4. Every minute
    1. Load command/answer from JSON files
    2. Send with with PingCommand
    3. Receive the PingAnswer from management server
    4. Remove the JSON file if state is COMPLETE/FAILED


Management server: (for reconcile command only)

  1. When create the reconcile command, insert a record with state CREATED into  reconcile_commands table
  2. When receive the PingCommand, Update command state and answer in reconcile_commands table
  3. When wait for the answer of command (reconcile command only)
    1. Every 10 seconds, check answer of the reconcile_commands table
    2. If answer is found, parse the answer
    3. Returna and continue with the answer
    4. No need to wait until timeout, if there are connection failures
  4. if timed out, update command state to TIMED_OUT in reconcile_commands table
  • Scenario 2: Agent crash (or force killing the agent process)
Agent interrupt the process or processing in backend

Management server

  1. When Host Status is determined as Down
  2. Update reconcile commands (to the agent) in reconcile_command table
    1. to INTERRUPTED state
  3. Reconcile the command
    1. see "3.6 Reconcile the command"
  • Scenario 3: Agent restart
Agent interrupt the process or processing in backend

Agent (when restart)

  1. When stop the agent, updates state of processes
    1. PROCESSING to INTERRUPTED
    2. PROCESSING_IN_BACKEND to DANGLED_IN_BACKEND
  2. When start the agent, updates state of processes 
    1. PROCESSING to INTERRUPTED
    2. PROCESSING_IN_BACKEND to DANGLED_IN_BACKEND
  3. Agent send CommandInfo to management server via PingCommand every minute
    1. with new state
  4. Management server reconcile commands every minute (from reconcile_commands table) , for commands in state
  • Scenario 4: Agent has completed but timed out
timed out on management server

Management server

  1. Update state_by_management to TIMED_OUT
  2. Update state_by_agent to COMPLETED (from PROCESSING)
  3. Reconcile the command
    1. see "3.6 Reconcile the command"
  4. TODO: what should be the correct resource state ?4
  • Scenario 5: Management restart

Agent process the command

  • has completed and send Answer to management server, but no action on management server
  • has not completed, and is processing command

Management server

  1. Update state_by_management to INTERRUPTED when mgmt server is stopped.
  2. When another management server is detected DOWN
    1. update reconcile_command to INTERRUPTED state by management server



3.5 State transitions (Agent)


Management server (all good): CREATED -> COMPLETED

...

            -> RECONCILED (all good) / RECONCILE_FAILED (failed, will retry) / RECONCILED_RETRY (success, but need more information, will retry)

Migrate VM

  • None

Migrate VM with volumes

...

  • if Running on dest, Running
  • If Paused on dest, Migrating

...

    • if source volume is found but destination is not found, mark source as Ready and dest as Destroy
    • if source volume is not found but destination is  found, mark source as Destroy and dest as Ready


See "4.3 Summary of test results" on how to reconcile the command


3.7 Limitations


  • Assumption
    • Each request does not have multiple commands with same name
    • for example, 700028267079401857-org.apache.cloudstack.storage.command.CopyCommand
  • Only support:
    • 3 commands
      • CopyCommand
      • MigrateCommand
      • MigrateVolumeCommand
    • resources in Migration state (vm, source volume, dest volume)
      • skipped if the resource state is not Migrating
      • skipped if the source or dest host/pool are inconsistent with command
    • Hypervisor: KVM
    • Storage: NFS, Local, Powerflex


3.8 To be Discussed (TODO)

  1. How to distribute the reconciiation tasks if there are multiple management servers ?
    1. now: Reconciliation is processed by first management server
  2. How to better handle the state DANGLED_IN_BACKEND ?
    1. now: no special action
  3. How to better handle the state COMPLETED and FAILED ? Do not reconcile via hosts if state_by_agent to COMPLETED ?
    1. now: no special action


4. Test cases


4.1 Backend commands of VM and volume migrations

Please note: 

Migration between NFS and Local requires the fix: https://github.com/apache/cloudstack/pull/10266


NFS to NFSNFS to LocalLocal to LocalLocal to NFSPowerflex to Powerflex (different envs)Powerflex <------> NFS

Migrate VM

(NFS or Powerflex)

PrepareForMigrationCommand (dest)

MigrateCommand (source)

---

PrepareForMigrationCommand (dest)

MigrateCommand (source)

-

Migrate VM with volumes


(NFS only)

CopyCommand (template to primary if needed)

CreateObjectCommand (new volume)

ModifyTargetsCommand

PrepareForMigrationCommand (dest)

MigrateCommand (source)

DeleteCommand (source)

same as "NFS to NFS"same as "NFS to NFS"same as "NFS to NFS"

Migrating a volume online with KVM from managed storage is not currently supported.

Pool [%s] is not compatible with volume [%s], skipping it.

Migrate ROOT Volume (of Running VM)


(Powerflex only)

KVM does not support volume live migrationdue to the limited possibility to refresh VM XML domain. Therefore, to live migrate a volume between storage pools, one must migrate the VM to a different host as well to force the VM XML domain update. Use 'migrateVirtualMachineWithVolumes' instead.samesamesame

MigrateVolumeCommand

  • Same volume ID
Storage pool pr503-t11980-kvm-ol8-kvm-pri3 is not suitable to migrate volume 

Migrate ROOT Volume (of Stopped VM)


(NFS or Powerflex)

CopyCommand (primary1 to secondary)

CopyCommand (secondary to primary2)

DeleteCommand (secondary)

DeleteCommand (primary1)

samesamesame

CopyCommand (primary1 to primary2)

  • Different volume IDs
same as above







Migrate DATA Volume (of Running VM

Migrate Volumes  (If volume is Migrating)

  • CopyCommand (from primary1 to secondary)
    • skipped.
    • change state from Creating to Destroy
    • clean volume_store_ref
  • CopyCommand (from secondary to primary2)
    • find an endpoint 
    • check if volume exists on primary2
    • check if volume size is changed on primary2
      • if yes, still Copying
      • if no, update state from Creating to Destroy, and clean volume_store_ref
  • CopyCommand (from primary1 to primary2)
    • check if volume exists on primary1
    • check if volume exists on primary2
    • check if volume size is changed on primary2
      • if yes, still Copying
      • if no, update state
  • MigrateVolumeCommand (from primary1 to primary2)
    • check if volume exists on primary1
    • check if volume exists on primary2
    • check if volume size is changed on primary2
      • if yes, still Migrating
      • if no, update state

3.7 Limitations

  • Assumption
    • Each request does not have multiple commands with same name
    • for example, 700028267079401857-org.apache.cloudstack.storage.command.CopyCommand
  • Only support:
    • 3 commands
      • CopyCommand
      • MigrateCommand
      • MigrateVolumeCommand
    • resources in Migration state (vm, source volume, dest volume)
      • skipped if the resource state is not Migrating
      • skipped if the source or dest host/pool are inconsistent with command
    • Hypervisor: KVM
    • Storage: NFS, Local, Powerflex

3.8 To be Discussed (TODO)

  1. How to distribute the reconciiation tasks if there are multiple management servers ?
    1. now: Reconciliation is processed by first management server
  2. How to better handle the state DANGLED_IN_BACKEND ?
    1. now: no special action
  3. How to better handle the state COMPLETED and FAILED ? Do not reconcile via hosts if state_by_agent to COMPLETED ?
    1. now: no special action

4. Test cases

4.1 Backend commands of some migrations

Please note: 

Migration between NFS and Local requires the fix: https://github.com/apache/cloudstack/pull/10266

Migrate VM with volumes

(NFS only)

NFS to NFSNFS to LocalLocal to LocalLocal to NFSPowerflex to PowerflexPowerflex <------> NFS

Migrate VM

(NFS or Powerflex)

PrepareForMigrationCommand (dest)

MigrateCommand (source)

---

PrepareForMigrationCommand (dest)

MigrateCommand (source)

-

CopyCommand (template to primary if needed)

CreateObjectCommand (new volume)

ModifyTargetsCommand

PrepareForMigrationCommand (dest)

MigrateCommand (source)

DeleteCommand (source)

same as "NFS to NFS"same as "NFS to NFS"same as "NFS to NFS"

Migrating a volume online with KVM from managed storage is not currently supported.

Pool [%s] is not compatible with volume [%s], skipping it.

Migrate ROOT Volume (of Running VM)

(Powerflex only)
KVM does not support volume live migrationdue to the limited possibility to refresh VM XML domain. Therefore, to live migrate a volume between storage pools, one must migrate the VM to a different host as well to force the VM XML domain update. Use 'migrateVirtualMachineWithVolumes' instead.samesamesame

MigrateVolumeCommand

  • Same volume
Storage pool pr503-t11980-kvm-ol8-kvm-pri3 is not suitable to migrate volume 
  • ID
same as above
Migrate DATA Migrate ROOT Volume (of Stopped VM)(NFS or Powerflex)

CopyCommand (primary1 to secondary)

CopyCommand (secondary to primary2)

DeleteCommand (secondary)

DeleteCommand (primary1)

samesamesame

CopyCommand (primary1 to primary2)

  • Different volume IDs
same as above
Migrate DATA Volume (of Running VMunattached)KVM does not support volume live migrationdue to the limited possibility to refresh VM XML domain. Therefore, to live migrate a volume between storage pools, one must migrate the VM to a different host as well to force the VM XML domain update. Use 'migrateVirtualMachineWithVolumes' instead.samesamesame

MigrateVolumeCommand

  • Same volume
same as above
Migrate DATA Volume (of Stopped VM)

CopyCommand (primary1 to secondary)

CopyCommand (secondary to primary2)

DeleteCommand (secondary)

DeleteCommand (primary1)

samesamesame

CopyCommand (primary1 to primary2)

  • Different volume IDs
same as above
Migrate DATA Volume (unattached)

CopyCommand (primary1 to secondary)

CopyCommand (secondary to primary2)

DeleteCommand (secondary)

DeleteCommand (primary1)

samesamesame

CopyCommand (primary1 to primary2)

  • Different volume IDs
same as above

4.2 How to test

...

  • systemctl restart cloudstack-agent

...

  • systemctl restart cloudstack-management

...

CopyCommand (primary1 to secondary)

CopyCommand (secondary to primary2)

DeleteCommand (secondary)

DeleteCommand (primary1)

samesamesame

CopyCommand (primary1 to primary2)

  • Different volume IDs
same as above


4.2 How to test

  • agent is restarted
    • systemctl restart cloudstack-agent
  • management server is restarted
    • systemctl restart cloudstack-management
  • agent and management server communication failure
    • iptables -I OUTPUT -p tcp -m tcp --dport 8250 -j DROP
    • iptables -D OUTPUT -p tcp -m tcp --dport 8250 -j DROP
  • Agent crash
    • pid=$(ps -ef |grep cloudstack-agent |grep -v grep |awk '{print $2}')
    • kill -9 $pid
  • Agent has completed but timed out


4.3 Summary of test results


...

Action

scenario 1: Connection failures between agent and management server

Scenario 4: Agent has completed but timed out

scenario 2

Agent crash (or force killing the agent process)

scenario 3

Agent is restarted manually

scenarios 5 (Needed ?)

mgmt server is restarted

How to test
  • iptables -I OUTPUT -p tcp -m tcp --dport 8250 -j DROP


  • iptables -D OUTPUT -p tcp -m tcp --dport 8250 -j DROP

...

  • pid=$(ps -ef |grep cloudstack-agent |grep -v grep |awk '{print $2}')

...

  • && kill -9 $pid

...

  • systemctl restart cloudstack-agent
  • systemctl restart cloudstack-management

1. Migrate VM

  • MigrateVMCommand
  • MigrateCommand (internal)


Supported by NFS or Powerflex

intermittent failure

    • agent updates answer when connection is back to normal (every reconcile.command.interval in PingRoutingCommand)
      • mgmt server saves the answer into DB
    • mgmt server gets the answer from DB every 10 seconds
    • mgmt server processes the answer if found
    • Operation times out


continuous failure

    • agent
      • Migration thread of VM [i-2-1198-VM] finished.
    •  management
      • Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance
    • Behavior:
      • Active migration command so scheduling a restart
      • Current: vm is Stopped on destination and Running on source (actually it is not if migration is successful)
      • New: vm is Running on destination host if found (via libvirt)
      • (new) check via destination host (if dest is Up)
        • if vm is Running on destination, consider the vm migration is successful.
  •  agent
    • agent is started again by systemctl
  • management server

4.3 Summary of test results

  • iptables -I OUTPUT -p tcp -m tcp --dport 8250 -j DROP
  • iptables -D OUTPUT -p tcp -m tcp --dport 8250 -j DROP
Action

scenario 1: Connection failures between agent and management server

Scenario 4: Agent has completed but timed out

scenario 2

Agent crash (or force killing the agent process)

scenario 3

Agent is restarted manually

scenarios 5 (Needed ?)

mgmt server is restarted

How to simulator
  • pid=$(ps -ef |grep cloudstack-agent |grep -v grep |awk '{print $2}') && kill -9 $pid
  • systemctl restart cloudstack-agent
  • systemctl restart cloudstack-management

Migrate VM

  • MigrateVMCommand
  • MigrateCommand (internal)

intermittent failure

    • agent updates answer when connection is back to normal (every ping.interval in PingRoutingCommand)
      • mgmt server saves the answer into DB
    • mgmt server gets the answer from DB every 10 seconds
    • mgmt server processes the answer if found
    • Operation times out

continuous failure

  • agent
    • Migration thread of VM [i-2-1198-VM] finished.
  •  managementNew:
    • Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance

  • Behavior:
  • Active migration command so scheduling a restart
  • Current: vm is Stopped on destination and Running on source (actually it is not if migration is successful)
    • (new) check via destination host (if dest is Up)
      • if vm is Running on destination, consider the vm migration is successful.
    • Current
      • VM is set to Stopped state on destination host and then Running on the source host
    • New:
          • vm is Running on destination
      • host
          • if found
      • (via libvirt)
      • (new) check via destination host (if dest is Up)
        • if vm is Running on destination, consider the vm migration is successful.
          • if not found, set to Stopped state on destination host and then Running on the source host (same as current)


      • Note:
          • In edge case, if the vm is migrated before agent is restarted (in 10 seconds) , the vm is Stopped (by ACS on destination host)
          • solved



    • agent 
      • when agent is restarted or disconnected
      • abort the migration job by a disconnect hook (MigrationCancelHook)
        • written by Marcus
        • tested with multiple mgmt servers
       agent
      • agent is started again by systemctl
    • management server
      • Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance
      • (new) check via destination host (if dest is Up)
        • if vm is Running on destination, consider the vm migration is successful.
      • Current: VM is set to Stopped state on destination host and then Running on the source host
      • New:
        • vm is Running on destination if found
        • if not found, set to Stopped state on destination host and then Running on the source host (same as current) - TODO


    • Note:
      • In edge case,
      • if
      • the vm is migrated before
      • agent is restarted (in 10 seconds)
      • aborting the job, the vm is Stopped (by ACS on destination host)
      • solved


    Issues

    • agent 
      • abort the migration job by disconnect hook (via libvirt)
        • written by Marcus
        • triggered when agent is restarted or disconnected
        • tested with multiple mgmt servers
    • management server
      • Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance
      • (new) check via destination host (if dest is Up)
        • if vm is Running on destination, consider the vm migration is successful.
      • Current: VM is set to Stopped state on destination host and then Running on the source host
      • New:
        • vm is Running on destination if found
        • if not found, set to Stopped state on destination host and then Running on the source host (same as current)
    • Note:
      • In edge case, the vm is migrated before aborting the job, the vm is Stopped (by ACS on destination host)
      • solved
    • mgmt server
      • Cancel left-over job-17783
      • Cleaning up Instance with Id: 1226
      • CleanUp Async Jobs after mgmt server maintenance (#8394)
      • vm is Running on source host
    • agent (new)
      • abort the migration job by disconnect hook
      • written by Marcus
    • Edge case: both scenario 1 and 5 happen
    • No issues found, ALL looks ok
    • mgmt server
      • Cancel left-over job-17783
      • Cleaning up Instance with Id: 1226
      • CleanUp Async Jobs after mgmt server maintenance (#8394)
      • vm is Running on source host (this will be updated when mgmt server gets vm report from agent)
    • agent (new)
      • abort the migration job by disconnect hook
      • written by Marcus
    • Edge case: both scenario 1 and 5 happen


    • No issues found, ALL looks ok


    2. Migrate VM with volumes

    • (between NFS)
    • MigrateVirtualMachine

    WithVolumeCmd

    • MigrateCommand (internal)


    Note: this action is not supported on PowerFlex


    Two volumes in DB

    • source volume
    • dest volume (last_id = source_volume_id)

    intermittent failure

      • same as above


    continuous failure

      • current: vm is Stopped on destination and Running on source  and pool (actually it is not if migration is successful)
      • New: vm is Running on destination host and pool if found on destination host
    • management server
      • Failed to migrate VM [VM instance xxx along with its volumes due to [com.cloud.utils.exception.CloudRuntimeException: Copy volume(s) to storage(s) xxx failed in StorageSystemDataMotionStrategy.copyAsync. Error message: [Commands 4446460207098233056 to Host 1 timed out after 21600].].
      • Current:
        • vm is Running on source host
        • volume is Ready on source pool
        • Bug: new volume is Migrating on destination pool (solved)
      • new
        • check via destination host (if dest is Up)
          • if vm is Running on destination, consider the vm migration is successful. 
        • (after fix) new volume is Destroy on destination pool if fail


    • Note:
      • In edge case, if the vm is

    Migrate VM with volumes

    • (between NFS)
    • MigrateVirtualMachine

    WithVolumeCmd

    • MigrateCommand (internal)

    Note: this action is not supported on PowerFlex

    intermittent failure

      • same as above

    continuous failure

      • current: vm is Stopped on destination and Running on source  and pool (actually it is not if migration is successful)
      • New: vm is Running on destination host and pool if found on destination host
    • management server
      • Failed to migrate VM [VM instance xxx along with its volumes due to [com.cloud.utils.exception.CloudRuntimeException: Copy volume(s) to storage(s) xxx failed in StorageSystemDataMotionStrategy.copyAsync. Error message: [Commands 4446460207098233056 to Host 1 timed out after 21600].].
      • Current:
        • vm is Running on source host
        • volume is Ready on source pool
        • new volume is Migrating on destination pool (solved)
      • new
        • check via destination host (if dest is Up)
          • if vm is Running on destination, consider the vm migration is successful. 
        • (after fix) new volume is Destroy on destination pool if fail
    • Note:
      • In edge case, if the vm is migrated before agent is restarted (in 10 seconds) , the vm is Stopped (by ACS on destination host)  
      • solved
    same as left

    Issues

    • management server
      • vm is Running on source host
      • Bug: both volumes are stuck at Migrating state (need to reconcile)
    • Edge case: both scenario 1 and 5 happen
      • VM is migrated but state is Running on source 


    How to reconcile (Verified):

    • only if If
        • VM is Running state
        and
        • volume is Migrating state
      • check via host (if source is Up)determine the state and update
      • if Running on dest, Running
      • If Paused on dest, Migrating
      • check disk path and update volume state
          • if source volume is found but destination is not found, mark source as Ready and dest as Destroy
          • if source volume is not found but destination is  found, mark source as Destroy and dest as Ready

    3. Migrate Volume (on NFS)

    • MigrateVolumeCmd (API)
    • CopyCommand (from primary1 to sec1)
    • CopyCommand (from sec1 to primary2)


    ROOT/DATA volume of Stopped VM

    intermittent failure

    • 1st CopyCommand: wait until connection is recoved
    • 2nd CopyCommand: wait until connection is recoved

    continuous failure

    • 1st CopyCommand: operation timeout
      • Resource [StoragePool:1] is unreachable: Volume [{"name":"ROOT-1159","uuid":"ed322109-8e54-47ca-8183-3c8f2a0f37e6"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 5302988561228766380 to Host 1 timed out after 1800].
      • Bug: new volume is Creating with empty path on destination pool (solved)
    • 2nd CopyCommand
      • Resource [StoragePool:1] is unreachable: Volume [{"name":"ROOT-1159","uuid":"ed322109-8e54-47ca-8183-3c8f2a0f37e6"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 7457960982925541434 to Host 1 timed out after 2400].
      • Bug: new volume is Creating with empty path on destination pool  (solved)


    New

    • new volume is Destroy state with empty path on destination pool


    • 1st CopyCommand
      • copy object failed: com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 15481123719087037 to Host 1 timed out after 1800
      • new volume is Destroy state
        • expungeVolumeAsync is called in copyVolumeCallBack


    • 2nd CopyCommand
      • copy failed com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 398850041998999773 to Host 1 timed out after 2400
      • new volume is Destroy state
        • expungeVolumeAsync is called in copyVolumeCallBack


    • No issues found, ALL looks good
    same as left

    Issues

    • 1st CopyCommand
      • new volume is Creating state
      • record in volume_store_ref is Creating state (to be cleaned)
    • 2nd CopyCommand
      • new volume is Creating state
      • record in volume_store_ref is Copying state (to be cleaned)


    How to reconcile (verified):

    • 1st CopyCommand
      • new volume is Destroy state and removed=now()
      • volume_store_ref with Creating state is removed
    • 2nd CopyCommand
      • check the size of new volume. If size is changed, the process is still ongoing.
      • if size is not changed, new volume is Destroy state and path is set (but not removed)
      • volume_store_ref with Copying state is removed

    4. Migrate Volume of Running VM (on Powerflex)

    • MigrateVolumeCommand

    TODO


    • agent: succeed
      • answer: {"volumePath":"1fe84f0700000002:vol-610-4d48-306f","result":true,"contextMap":{},"wait":0,"bypassHostMaintenance":false}
    • management servermanagement
      • Resource [StoragePool:12] is unreachable: Migrate volume failed: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 2520889891420635168 to Host 1 timed out after 2400
      • Failed to migrate PowerFlex volume
      • volume path is set to new, then set to old when fail (looks good)
        • revertBlockCopyVolumeOperations
    • No issues found, ALL looks good

    same as left

    TODO

    Migrate Volume of Stopped VM (on Powerflex)

    • CopyCommand
      • Bug: Actually, on KVM host, the VM is running with new volume, but the new volume is removed
        • Input/Output error, read-only system
      • New:
        • check volume statistics on destination pool via ScaleIO gateway
        • if allocation size = provisioned size, volume migration is done.  remove volume on source pool. 
        • if allocation size is not same to provisioned size, volume migration failes, remove volume on destination pool. 

    intermittent failure

    • CopyCommand: wait until connection is recoved
    continuous failure
    • management
      • Resource [StoragePool:
      11
      • 12] is unreachable:
      Volume [{"name":"ROOT-533","uuid":"ff8a53a0-d513-4bfc-8920-352948ea066b"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: Failed to
      • Migrate volume failed: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands
      8886446489732120778
      • 2520889891420635168 to Host 1 timed out after 2400
      ]
      • 2400 = wait (1200) * 2 
      • new volume is Expunged
        • due to expungeVolumeAsync in copyManagedVolumeCallBack
      • Not an issue since the VM is Stopped so no data loss
    • management server
      • Resource [StoragePool:12] is unreachable: Volume [{"name":"ROOT-531","uuid":"4225516d-52fe-465d-a46d-81f97a2e55b2"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 3050062847636668430 to Host 1 timed out after 2400].
      • new volume is Expunged
        • due to expungeVolumeAsync in copyManagedVolumeCallBack
    • No issues found, ALL looks good

    Exception while reconcile : TO be fixed

      • Failed to migrate PowerFlex volume
      • volume path is set to new, then set to old when fail (looks good)
        • revertBlockCopyVolumeOperations
    • No issues found, ALL looks good


    changes on agent 

    • it uses "dm.blockCopy" (introduced by Hari)
    • when agent is disconnected or restarted
      • abort the block copy job by a disconnect hook (VolumeMigrationCancelHook)
      • similar as MigrationCancelHook

    same as left

    Issue:

    • one volume in Migrating in DB
      • path, iscsi_name, pool_id are info on destination pool 
    • when mgmt server is restarted, it updates all Migrating volume to Ready
      • this is disabled if reconcile is enabled

    How to reconcile:

    • only if vm is Running and volume in Migrating state
      • now: 
    • The current pool is destination pool of reconcile command
    • check vm state and vm disks via agent
      • if volume is Ready on source pool,
        • revert path, iscsi_name and pool_id to source pool
        • create a dummy volume on dest pool
        • remove the dummy volume
      • if  volume is not found on source pool
        • update state to Ready
        • create a dummy volume on source pool
        • remove the dummy volume
      • if vm is running on source pool (check by fullpath) = volume is ready on source pool
      • if vm is running on dest pool, update volume state to Ready, and remove volume from source pool

    5. Migrate Volume of Stopped VM (on Powerflex)

    • CopyCommand


    intermittent failure

    • CopyCommand: wait until connection is recoved

    continuous failure

    • Resource [StoragePool:11] is unreachable: Volume [{"name":"ROOT-533","uuid":"ff8a53a0-d513-4bfc-8920-352948ea066b"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 8886446489732120778 to Host 1 timed out after 2400]
      • 2400 = wait (1200) * 2 
      • new volume is Expunged
        • due to expungeVolumeAsync in copyManagedVolumeCallBack
      • Not an issue since the VM is Stopped so no data loss
    • management server
      • Resource [StoragePool:12] is unreachable: Volume [{"name":"ROOT-531","uuid":"4225516d-52fe-465d-a46d-81f97a2e55b2"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 3050062847636668430 to Host 1 timed out after 2400].
      • new volume is Expunged
        • due to expungeVolumeAsync in copyManagedVolumeCallBack


    • No issues found, ALL looks good

    same as left

    Issue: two volumes in DB

    • source volume: Migrating state
    • dest volume: Creating state
      • dest.last_id = source_volume_id

    How to reconcile

    • check volume information via host
    • determine state of source and dest volumes
    • if source volume is Ready
      • update source from Migrating to Ready
      • update dest volume from Creating to Destroy
    • if dest volume is Ready
      • update dest from Creating to Ready
      • update source volume from Migrating to Destroy


    TO be fixed

    • when migrate volume from powerflex to powerflex with same system id, it will not create CopyCommand

    same as left

    TODO