Versions Compared

Key

  • This line was added.
  • This line was removed.
  • Formatting was changed.

...

            -> RECONCILED (all good) / RECONCILE_FAILED (failed, will retry) / RECONCILED_RETRY (success, but need more information, will retry)


3.7 Summary


Action

scenario 1

Connection failures between

the

agent and

the

management

server

scenario 2

Agent crash (or force killing the agent process)

scenario 3

Agent is restarted manually

scenarios 5 (Needed ?)

mgmt

server

is restartedMigrate VMagent

  • iptables -I OUTPUT -p tcp -m tcp --dport 8250 -j DROP
    • iptables -D OUTPUT -p tcp -m tcp --dport 8250 -j DROP
  • intermittent failure
    • agent updates answer when connection is back to normal
    • mgmt server gets the answer from DB every 10 seconds
    • mgmt server processes the answer if found
    • Operation times out
continuous failure

scenario 2

Agent crash (or force killing the agent process)

  • pid=$(ps -ef |grep cloudstack-agent |grep -v grep |awk '{print $2}') && kill -9 $pid

scenario 3

Agent is restarted manually

  • systemctl restart cloudstack-agent

scenarios 5 (Needed ?)

mgmt server is restarted

  • systemctl restart cloudstack-management

Migrate VM

  • MigrateVMCommand
  • MigrateCommand (internal)

intermittent failure

    • agent updates answer when connection is back to normal
    • mgmt server gets the answer from DB every 10 seconds
    • mgmt server processes the answer if found
    • Operation times out


continuous failure

    • agent
        • Migration thread of VM [i-2-1198-VM] finished.
      •  management
        • Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance
      • Behavior:
        • Active migration command so scheduling a restart
        • Current: vm is Stopped on destination and Running on source (actually it is not if migration is successful)
        • New: vm is Running on destination host if found
        • (new) check via destination host (if dest is Up)
          • if vm is Running on destination, consider the vm migration is successful.
    kill -9  $(ps -ef |grep cloudstack-agent |grep -v grep |awk '{print $2}')
    •  agent
      • agent is started again by systemctl
    • management server

      • Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance

      • (new) check via destination host (if dest is Up)
        • if vm is Running on destination, consider the vm migration is successful.
      • Current: VM is Running on source host
      • New: vm is Running on destination if found


    • Note:
        • In edge case, if the vm is migrated before agent is restarted (in 10 seconds) , the vm is Stopped (by ACS on destination host)
        • solved
    • systemctl restart cloudstack-agent

    Migrate Volume (on NFS)

    • CopyCommand (from primary1 to sec1)
    • CopyCommand (from sec1 to primary2)

    ROOT/DATA volume of Stopped VM

    ROOT/DATA volume of Stopped VM



    • agent 
      • abort the migration job by disconnect hook
      • written by Marcus
    • management server
      • Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance
      • (new) check via destination host (if dest is Up)
        • if vm is Running on destination, consider the vm migration is successful.
      • Current: VM is Running on source host
      • New: vm is Running on destination if found


    • Note:
      • In edge case, the vm is migrated before aborting the job, the vm is Stopped (by ACS on destination host)


    If VM is Migrating (TO be verified)

    • check via source host (if source is Up)
    • check via destination host (if dest is Up)
    • determine the state and update
      • if Running on dest, Running
      • If Paused on dest, Migrating
      • If Stopped on dest, Running (if found on source) or Stopped (if not found on source)

    Migrate VM with volumes

    • (between NFS)
    • MigrateVirtualMachineWithVolumeCmd
    • MigrateCommand (internal)


    Note: this action is not supported on PowerFlex

    intermittent failure

      • same as above


    continuous failure

      • current: vm is Stopped on destination and Running on source  and pool (actually it is not if migration is successful)
      • New: vm is Running on destination host and pool if found on destination host
    • management server
      • Failed to migrate VM [VM instance xxx along with its volumes due to [com.cloud.utils.exception.CloudRuntimeException: Copy volume(s) to storage(s) xxx failed in StorageSystemDataMotionStrategy.copyAsync. Error message: [Commands 4446460207098233056 to Host 1 timed out after 21600].].
      • Current:
        • vm is Running on source host
        • volume is Ready on source pool
        • new volume is Migrating on destination pool (solved)
      • new
        • check via destination host (if dest is Up)
          • if vm is Running on destination, consider the vm migration is successful. 
        • (after fix) new volume is Destroy on destination pool if fail


    • Note:
      • In edge case, if
    the vm is migrated before agent is restarted (in 10 seconds) , the vm is Stopped (by ACS on destination host)
  • solved
  • same as left
      • the vm is migrated before agent is restarted (in 10 seconds) , the vm is Stopped (by ACS on destination host)  
      • solved
    same as left

    Migrate Volume (on NFS)

    • MigrateVolumeCmd (API)
    • CopyCommand (from primary1 to sec1)
    • CopyCommand (from sec1 to primary2)


    ROOT/DATA volume of Stopped VM

    intermittent failure

    • 1st CopyCommand: wait until connection is recoved
    • 2nd CopyCommand: wait until connection is recoved

    continuous failure

    • 1st CopyCommand: operation timeout
      • Resource [StoragePool:1] is unreachable: Volume [{"name":"ROOT-1159","uuid":"ed322109-8e54-47ca-8183-3c8f2a0f37e6"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 5302988561228766380 to Host 1 timed out after 1800].
      • new volume is Creating with empty path on destination pool (solved)
    • 2nd CopyCommand
      • Resource [StoragePool:1] is unreachable: Volume [{"name":"ROOT-1159","uuid":"ed322109-8e54-47ca-8183-3c8f2a0f37e6"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 7457960982925541434 to Host 1 timed out after 2400].
      • new volume is Creating with empty path on destination pool  (solved)


    New

    • new volume is Destroy state with empty path on destination pool


    • 1st CopyCommand
      • copy object failed: com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 15481123719087037 to Host 1 timed out after 1800
      • new volume is Destroy state
        • expungeVolumeAsync is called in copyVolumeCallBack


    • 2nd CopyCommand
      • copy failed com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 398850041998999773 to Host 1 timed out after 2400
      • new volume is Destroy state
        • expungeVolumeAsync is called in copyVolumeCallBack


    • ALL looks good
    same as left

    Migrate Volume (on Powerflex)

    • Running VM
    • Stopped VM

    Current:

    TODO

    Current:

    TODO

    Current:

    TODO



    Migrate VM with volumes  (If VM is Migrating)

    ...