Versions Compared

Key

  • This line was added.
  • This line was removed.
  • Formatting was changed.

...

scenario 2

Agent crash (or force killing the agent process)scenario 3

Agent is restarted manually

scenarios 5 (Needed ?)

mgmt server is restarted
Action

scenario 1: Connection failures between agent and management server

Scenario 4: Agent has completed but timed out

scenario 2

Agent crash (or force killing the agent process)

scenario 3

Agent is restarted manually

scenarios 5 (Needed ?)

mgmt server is restarted

How to simulator
  • iptables -I OUTPUT -p tcp -m tcp --dport 8250 -j DROP


  • iptables -D OUTPUT -p tcp -m tcp --dport 8250 -j DROP
  • pid=$(ps -ef |grep cloudstack-agent |grep -v grep |awk '{print $2}') && kill -9 $pid
  • systemctl restart cloudstack-agent
  • systemctl restart cloudstack-management

Migrate VM

  • MigrateVMCommand
  • MigrateCommand (internal)

intermittent failure

    • agent updates answer when connection is back to normal (every ping.interval in PingRoutingCommand)
      • mgmt server saves the answer into DB
    • mgmt server gets the answer from DB every 10 seconds
    • mgmt server processes the answer if found
    • Operation times out


continuous failure

    • agent
      • Migration thread of VM [i-2-1198-VM] finished.
    •  management
      • Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance
    • Behavior:
      • Active migration command so scheduling a restart
      • Current: vm is Stopped on destination and Running on source (actually it is not if migration is successful)
      • New: vm is Running on destination host if found (via libvirt)
      • (new) check via destination host (if dest is Up)
        • if vm is Running on destination, consider the vm migration is successful.
  •  agent
    • agent is started again by systemctl
  • management server

    • Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance

    • (new) check via destination host (if dest is Up)
      • if vm is Running on destination, consider the vm migration is successful.
    • Current: VM is set to Stopped state on destination host and then Running on the source host
    • New:
      • vm is Running on destination if found
      • if not found, set to Stopped state on destination host and then Running on the source host (same as current) - TODO


  • Note:
      • In edge case, if the vm is migrated before agent is restarted (in 10 seconds) , the vm is Stopped (by ACS on destination host)
      • solved



  • agent 
    • abort the migration job by disconnect hook (via libvirt)
      • written by Marcus
      • triggered when agent is restarted or disconnected
      • tested with multiple mgmt servers
  • management server
    • Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance
    • (new) check via destination host (if dest is Up)
      • if vm is Running on destination, consider the vm migration is successful.
    • Current: VM is set to Stopped state on destination host and then Running on the source host
    • New:
      • vm is Running on destination if found
      • if not found, set to Stopped state on destination host and then Running on the source host (same as current)


  • Note:
    • In edge case, the vm is migrated before aborting the job, the vm is Stopped (by ACS on destination host)
    • solved


  • mgmt server
    • Cancel left-over job-17783
    • Cleaning up Instance with Id: 1226
    • CleanUp Async Jobs after mgmt server maintenance (#8394)
    • vm is Running on source host
  • agent (new)
    • abort the migration job by disconnect hook
    • written by Marcus
  • Edge case: both scenario 1 and 5 happen


  • No issues found, ALL looks ok


Migrate VM with volumes

  • (between NFS)
  • MigrateVirtualMachine

WithVolumeCmd

  • MigrateCommand (internal)


Note: this action is not supported on PowerFlex

intermittent failure

    • same as above


continuous failure

    • current: vm is Stopped on destination and Running on source  and pool (actually it is not if migration is successful)
    • New: vm is Running on destination host and pool if found on destination host
  • management server
    • Failed to migrate VM [VM instance xxx along with its volumes due to [com.cloud.utils.exception.CloudRuntimeException: Copy volume(s) to storage(s) xxx failed in StorageSystemDataMotionStrategy.copyAsync. Error message: [Commands 4446460207098233056 to Host 1 timed out after 21600].].
    • Current:
      • vm is Running on source host
      • volume is Ready on source pool
      • new volume is Migrating on destination pool (solved)
    • new
      • check via destination host (if dest is Up)
        • if vm is Running on destination, consider the vm migration is successful. 
      • (after fix) new volume is Destroy on destination pool if fail


  • Note:
    • In edge case, if the vm is migrated before agent is restarted (in 10 seconds) , the vm is Stopped (by ACS on destination host)  
    • solved
same as left
  • management server
    • vm is Running on source host
    • both volumes are stuck at Migrating state (need to reconcile)
  • Edge case: both scenario 1 and 5 happen
    • VM is migrated but state is Running on source 


How to reconcile (Verified):

  • If VM is Running state and volume is Migrating state
  • check via host (if source is Up)
  • determine the state and update
    • if Running on dest, Running
    • If Paused on dest, Migrating
  • check disk path and update volume state
      • if source volume is found but destination is not found, mark source as Ready and dest as Destroy
      • if source volume is not found but destination is  found, mark source as Destroy and dest as Ready

Migrate Volume (on NFS)

  • MigrateVolumeCmd (API)
  • CopyCommand (from primary1 to sec1)
  • CopyCommand (from sec1 to primary2)


ROOT/DATA volume of Stopped VM

intermittent failure

  • 1st CopyCommand: wait until connection is recoved
  • 2nd CopyCommand: wait until connection is recoved

continuous failure

  • 1st CopyCommand: operation timeout
    • Resource [StoragePool:1] is unreachable: Volume [{"name":"ROOT-1159","uuid":"ed322109-8e54-47ca-8183-3c8f2a0f37e6"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 5302988561228766380 to Host 1 timed out after 1800].
    • new volume is Creating with empty path on destination pool (solved)
  • 2nd CopyCommand
    • Resource [StoragePool:1] is unreachable: Volume [{"name":"ROOT-1159","uuid":"ed322109-8e54-47ca-8183-3c8f2a0f37e6"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 7457960982925541434 to Host 1 timed out after 2400].
    • new volume is Creating with empty path on destination pool  (solved)


New

  • new volume is Destroy state with empty path on destination pool


  • 1st CopyCommand
    • copy object failed: com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 15481123719087037 to Host 1 timed out after 1800
    • new volume is Destroy state
      • expungeVolumeAsync is called in copyVolumeCallBack


  • 2nd CopyCommand
    • copy failed com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 398850041998999773 to Host 1 timed out after 2400
    • new volume is Destroy state
      • expungeVolumeAsync is called in copyVolumeCallBack


  • No issues found, ALL looks good
same as left
  • 1st CopyCommand
    • new volume is Creating state
    • record in volume_store_ref is Creating state (to be cleaned)
  • 2nd CopyCommand
    • new volume is Creating state
    • record in volume_store_ref is Copying state (to be cleaned)


How to reconcile (verified):

  • 1st CopyCommand
    • new volume is Destroy state and removed=now()
    • volume_store_ref with Creating state is removed
  • 2nd CopyCommand
    • check the size of new volume. If size is changed, the process is still ongoing.
    • if size is not changed, new volume is Destroy state and path is set (but not removed)
    • volume_store_ref with Copying state is removed

Migrate Volume of Running VM (on Powerflex)

  • MigrateVolumeCommand


TODO


TODO


TODO


TODO

Migrate Volume of Stopped VM (on Powerflex)

  • CopyCommand


TODO


TODO


TODO


TODO

...