| Action | scenario 1: Connection failures between agent and management server Scenario 4: Agent has completed but timed out - iptables -I OUTPUT -p tcp -m tcp --dport 8250 -j DROP
- iptables -D OUTPUT -p tcp -m tcp --dport 8250 -j DROP
| scenario 2 Agent crash (or force killing the agent process) - pid=$(ps -ef |grep cloudstack-agent |grep -v grep |awk '{print $2}') && kill -9 $pid
| scenario 3 Agent is restarted manually - systemctl restart cloudstack-agent
| scenarios 5 (Needed ?) mgmt server is restarted - systemctl restart cloudstack-management
|
|---|
Migrate VM - MigrateVMCommand
- MigrateCommand (internal)
| intermittent failure - agent updates answer when connection is back to normal
- mgmt server gets the answer from DB every 10 seconds
- mgmt server processes the answer if found
- Operation times out
continuous failure - agent
- Migration thread of VM [i-2-1198-VM] finished.
- management
- Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance
- Behavior:
- Active migration command so scheduling a restart
- Current: vm is Stopped on destination and Running on source (actually it is not if migration is successful)
- New: vm is Running on destination host if found
- (new) check via destination host (if dest is Up)
- if vm is Running on destination, consider the vm migration is successful.
| - agent
- agent is started again by systemctl
- Note:
- In edge case, if the vm is migrated before agent is restarted (in 10 seconds) , the vm is Stopped (by ACS on destination host)
- solved
| - agent
- abort the migration job by disconnect hook
- written by Marcus
- management server
- Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance
- (new) check via destination host (if dest is Up)
- if vm is Running on destination, consider the vm migration is successful.
- Current: VM is Running on source host
- New: vm is Running on destination if found
- Note:
- In edge case, the vm is migrated before aborting the job, the vm is Stopped (by ACS on destination host)
- solved
| - mgmt server
- Cancel left-over job-17783
- Cleaning up Instance with Id: 1226
- CleanUp Async Jobs after mgmt server maintenance (#8394)
- vm is Running on source host
- agent (new)
- abort the migration job by disconnect hook
- written by Marcus
Edge case: both scenario 1 and 5 happen - VM is migrated but state is Running on source
|
Migrate VM with volumes - (between NFS)
- MigrateVirtualMachine
WithVolumeCmd - MigrateCommand (internal)
Note: this action is not supported on PowerFlex | intermittent failure
continuous failure - current: vm is Stopped on destination and Running on source and pool (actually it is not if migration is successful)
- New: vm is Running on destination host and pool if found on destination host
| - management server
- Failed to migrate VM [VM instance xxx along with its volumes due to [com.cloud.utils.exception.CloudRuntimeException: Copy volume(s) to storage(s) xxx failed in StorageSystemDataMotionStrategy.copyAsync. Error message: [Commands 4446460207098233056 to Host 1 timed out after 21600].].
- Current:
- vm is Running on source host
- volume is Ready on source pool
- new volume is Migrating on destination pool (solved)
- new
- check via destination host (if dest is Up)
- if vm is Running on destination, consider the vm migration is successful.
- (after fix) new volume is Destroy on destination pool if fail
- Note:
- In edge case, if the vm is migrated before agent is restarted (in 10 seconds) , the vm is Stopped (by ACS on destination host)
- solved
| same as left | - management server
- vm is Running on source host
- new volume is stuck at Migrating state (need to reconcile)
if Volume is Migrating, reconcile
|
Migrate Volume (on NFS) - CopyCommand (from primary1 to sec1)
- CopyCommand (from sec1 to primary2)
ROOT/DATA volume of Stopped VM | intermittent failure - 1st CopyCommand: wait until connection is recoved
- 2nd CopyCommand: wait until connection is recoved
continuous failure - 1st CopyCommand: operation timeout
- Resource [StoragePool:1] is unreachable: Volume [{"name":"ROOT-1159","uuid":"ed322109-8e54-47ca-8183-3c8f2a0f37e6"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 5302988561228766380 to Host 1 timed out after 1800].
- new volume is Creating with empty path on destination pool (solved)
- 2nd CopyCommand
- Resource [StoragePool:1] is unreachable: Volume [{"name":"ROOT-1159","uuid":"ed322109-8e54-47ca-8183-3c8f2a0f37e6"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 7457960982925541434 to Host 1 timed out after 2400].
- new volume is Creating with empty path on destination pool (solved)
New - new volume is Destroy state with empty path on destination pool
| - 1st CopyCommand
- copy object failed: com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 15481123719087037 to Host 1 timed out after 1800
- new volume is Destroy state
- expungeVolumeAsync is called in copyVolumeCallBack
- 2nd CopyCommand
- copy failed com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 398850041998999773 to Host 1 timed out after 2400
- new volume is Destroy state
- expungeVolumeAsync is called in copyVolumeCallBack
| same as left |
|
Migrate Volume (on Powerflex) | Current: TODO | Current: TODO | Current: TODO |
|