...
-> RECONCILED (all good) / RECONCILE_FAILED (failed, will retry) / RECONCILED_RETRY (success, but need more information, will retry)
3.7 Summary
| Action | scenario 1 Connection failures between |
|---|
the the serverscenario 2 Agent crash (or force killing the agent process) | scenario 3 Agent is restarted manually | scenarios 5 (Needed ?)
mgmt is restartedMigrate VMagent- iptables -I OUTPUT -p tcp -m tcp --dport 8250 -j DROP
- iptables -D OUTPUT -p tcp -m tcp --dport 8250 -j DROP
|
|---|
- intermittent failure
- agent updates answer when connection is back to normal
- mgmt server gets the answer from DB every 10 seconds
- mgmt server processes the answer if found
- Operation times out
continuous failurescenario 2 Agent crash (or force killing the agent process) - pid=$(ps -ef |grep cloudstack-agent |grep -v grep |awk '{print $2}') && kill -9 $pid
| scenario 3 Agent is restarted manually - systemctl restart cloudstack-agent
| scenarios 5 (Needed ?) mgmt server is restarted - systemctl restart cloudstack-management
|
|---|
Migrate VM - MigrateVMCommand
- MigrateCommand (internal)
| intermittent failure - agent updates answer when connection is back to normal
- mgmt server gets the answer from DB every 10 seconds
- mgmt server processes the answer if found
- Operation times out
continuous failure |
- Migration thread of VM [i-2-1198-VM] finished.
- management
- Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance
- Behavior:
- Active migration command so scheduling a restart
- Current: vm is Stopped on destination and Running on source (actually it is not if migration is successful)
- New: vm is Running on destination host if found
- (new) check via destination host (if dest is Up)
- if vm is Running on destination, consider the vm migration is successful.
|
kill -9 $(ps -ef |grep cloudstack-agent |grep -v grep |awk '{print $2}') | - agent
- agent is started again by systemctl
- Note:
- In edge case, if the vm is migrated before agent is restarted (in 10 seconds) , the vm is Stopped (by ACS on destination host)
- solved
|
- systemctl restart cloudstack-agent
Migrate Volume (on NFS)
- CopyCommand (from primary1 to sec1)
- CopyCommand (from sec1 to primary2)
ROOT/DATA volume of Stopped VM
ROOT/DATA volume of Stopped VM
| - agent
- abort the migration job by disconnect hook
- written by Marcus
- management server
- Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance
- (new) check via destination host (if dest is Up)
- if vm is Running on destination, consider the vm migration is successful.
- Current: VM is Running on source host
- New: vm is Running on destination if found
- Note:
- In edge case, the vm is migrated before aborting the job, the vm is Stopped (by ACS on destination host)
| If VM is Migrating (TO be verified) - check via source host (if source is Up)
- check via destination host (if dest is Up)
- determine the state and update
- if Running on dest, Running
- If Paused on dest, Migrating
- If Stopped on dest, Running (if found on source) or Stopped (if not found on source)
|
Migrate VM with volumes - (between NFS)
- MigrateVirtualMachineWithVolumeCmd
- MigrateCommand (internal)
Note: this action is not supported on PowerFlex | intermittent failure
continuous failure - current: vm is Stopped on destination and Running on source and pool (actually it is not if migration is successful)
- New: vm is Running on destination host and pool if found on destination host
| - management server
- Failed to migrate VM [VM instance xxx along with its volumes due to [com.cloud.utils.exception.CloudRuntimeException: Copy volume(s) to storage(s) xxx failed in StorageSystemDataMotionStrategy.copyAsync. Error message: [Commands 4446460207098233056 to Host 1 timed out after 21600].].
- Current:
- vm is Running on source host
- volume is Ready on source pool
- new volume is Migrating on destination pool (solved)
- new
- check via destination host (if dest is Up)
- if vm is Running on destination, consider the vm migration is successful.
- (after fix) new volume is Destroy on destination pool if fail
|
the vm is migrated before agent is restarted (in 10 seconds) , the vm is Stopped (by ACS on destination host)solvedsame as left | - the vm is migrated before agent is restarted (in 10 seconds) , the vm is Stopped (by ACS on destination host)
- solved
| same as left |
|
Migrate Volume (on NFS) - CopyCommand (from primary1 to sec1)
- CopyCommand (from sec1 to primary2)
ROOT/DATA volume of Stopped VM | intermittent failure - 1st CopyCommand: wait until connection is recoved
- 2nd CopyCommand: wait until connection is recoved
continuous failure - 1st CopyCommand: operation timeout
- Resource [StoragePool:1] is unreachable: Volume [{"name":"ROOT-1159","uuid":"ed322109-8e54-47ca-8183-3c8f2a0f37e6"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 5302988561228766380 to Host 1 timed out after 1800].
- new volume is Creating with empty path on destination pool (solved)
- 2nd CopyCommand
- Resource [StoragePool:1] is unreachable: Volume [{"name":"ROOT-1159","uuid":"ed322109-8e54-47ca-8183-3c8f2a0f37e6"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 7457960982925541434 to Host 1 timed out after 2400].
- new volume is Creating with empty path on destination pool (solved)
New - new volume is Destroy state with empty path on destination pool
| - 1st CopyCommand
- copy object failed: com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 15481123719087037 to Host 1 timed out after 1800
- new volume is Destroy state
- expungeVolumeAsync is called in copyVolumeCallBack
- 2nd CopyCommand
- copy failed com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 398850041998999773 to Host 1 timed out after 2400
- new volume is Destroy state
- expungeVolumeAsync is called in copyVolumeCallBack
| same as left |
|
Migrate Volume (on Powerflex) | Current: TODO | Current: TODO | Current: TODO |
|
Migrate VM with volumes (If VM is Migrating)
...