Versions Compared

Key

  • This line was added.
  • This line was removed.
  • Formatting was changed.

...

4.3 Summary of test results


Action

scenario 1: Connection failures between agent and management server

Scenario 4: Agent has completed but timed out

  • iptables -I OUTPUT -p tcp -m tcp --dport 8250 -j DROP
    • iptables -D OUTPUT -p tcp -m tcp --dport 8250 -j DROP

scenario 2

Agent crash (or force killing the agent process)

  • pid=$(ps -ef |grep cloudstack-agent |grep -v grep |awk '{print $2}') && kill -9 $pid

scenario 3

Agent is restarted manually

  • systemctl restart cloudstack-agent

scenarios 5 (Needed ?)

mgmt server is restarted

  • systemctl restart cloudstack-management

Migrate VM

  • MigrateVMCommand
  • MigrateCommand (internal)

intermittent failure

    • agent updates answer when connection is back to normal
    • mgmt server gets the answer from DB every 10 seconds
    • mgmt server processes the answer if found
    • Operation times out


continuous failure

    • agent
      • Migration thread of VM [i-2-1198-VM] finished.
    •  management
      • Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance
    • Behavior:
      • Active migration command so scheduling a restart
      • Current: vm is Stopped on destination and Running on source (actually it is not if migration is successful)
      • New: vm is Running on destination host if found
      • (new) check via destination host (if dest is Up)
        • if vm is Running on destination, consider the vm migration is successful.
  •  agent
    • agent is started again by systemctl
  • management server

    • Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance

    • (new) check via destination host (if dest is Up)
      • if vm is Running on destination, consider the vm migration is successful.
    • Current: VM is Running on source host
    • New: vm is Running on destination if found


  • Note:
      • In edge case, if the vm is migrated before agent is restarted (in 10 seconds) , the vm is Stopped (by ACS on destination host)
      • solved



  • agent 
    • abort the migration job by disconnect hook
    • written by Marcus
  • management server
    • Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance
    • (new) check via destination host (if dest is Up)
      • if vm is Running on destination, consider the vm migration is successful.
    • Current: VM is Running on source host
    • New: vm is Running on destination if found


  • Note:
    • In edge case, the vm is migrated before aborting the job, the vm is Stopped (by ACS on destination host)
    • solved


  • mgmt server
    • Cancel left-over job-17783
    • Cleaning up Instance with Id: 1226
    • CleanUp Async Jobs after mgmt server maintenance (#8394)
    • vm is Running on source host
  • agent (new)
    • abort the migration job by disconnect hook
    • written by Marcus


Edge case: both scenario 1 and 5 happen

  • VM is migrated but state is Running on source

Migrate VM with volumes

  • (between NFS)
  • MigrateVirtualMachine

WithVolumeCmd

  • MigrateCommand (internal)


Note: this action is not supported on PowerFlex

intermittent failure

    • same as above


continuous failure

    • current: vm is Stopped on destination and Running on source  and pool (actually it is not if migration is successful)
    • New: vm is Running on destination host and pool if found on destination host
  • management server
    • Failed to migrate VM [VM instance xxx along with its volumes due to [com.cloud.utils.exception.CloudRuntimeException: Copy volume(s) to storage(s) xxx failed in StorageSystemDataMotionStrategy.copyAsync. Error message: [Commands 4446460207098233056 to Host 1 timed out after 21600].].
    • Current:
      • vm is Running on source host
      • volume is Ready on source pool
      • new volume is Migrating on destination pool (solved)
    • new
      • check via destination host (if dest is Up)
        • if vm is Running on destination, consider the vm migration is successful. 
      • (after fix) new volume is Destroy on destination pool if fail


  • Note:
    • In edge case, if the vm is migrated before agent is restarted (in 10 seconds) , the vm is Stopped (by ACS on destination host)  
    • solved
same as left
  • management server
    • vm is Running on source host
    • new volume is stuck at Migrating state (need to reconcile)

if Volume is Migrating, reconcile


Migrate Volume (on NFS)

  • MigrateVolumeCmd (API)
  • CopyCommand (from primary1 to sec1)
  • CopyCommand (from sec1 to primary2)


ROOT/DATA volume of Stopped VM

intermittent failure

  • 1st CopyCommand: wait until connection is recoved
  • 2nd CopyCommand: wait until connection is recoved

continuous failure

  • 1st CopyCommand: operation timeout
    • Resource [StoragePool:1] is unreachable: Volume [{"name":"ROOT-1159","uuid":"ed322109-8e54-47ca-8183-3c8f2a0f37e6"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 5302988561228766380 to Host 1 timed out after 1800].
    • new volume is Creating with empty path on destination pool (solved)
  • 2nd CopyCommand
    • Resource [StoragePool:1] is unreachable: Volume [{"name":"ROOT-1159","uuid":"ed322109-8e54-47ca-8183-3c8f2a0f37e6"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 7457960982925541434 to Host 1 timed out after 2400].
    • new volume is Creating with empty path on destination pool  (solved)


New

  • new volume is Destroy state with empty path on destination pool


  • 1st CopyCommand
    • copy object failed: com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 15481123719087037 to Host 1 timed out after 1800
    • new volume is Destroy state
      • expungeVolumeAsync is called in copyVolumeCallBack


  • 2nd CopyCommand
    • copy failed com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 398850041998999773 to Host 1 timed out after 2400
    • new volume is Destroy state
      • expungeVolumeAsync is called in copyVolumeCallBack


  • ALL looks good
same as left

Migrate Volume (on Powerflex)

  • Running VM
  • Stopped VM

Current:

TODO

Current:

TODO

Current:

TODO