| Action | scenario 1: Connection failures between agent and management server Scenario 4: Agent has completed but timed out | scenario 2 Agent crash (or force killing the agent process) | scenario 3 Agent is restarted manually | scenarios 5 (Needed ?) mgmt server is restarted |
|---|
| How to simulator | - iptables -I OUTPUT -p tcp -m tcp --dport 8250 -j DROP
- iptables -D OUTPUT -p tcp -m tcp --dport 8250 -j DROP
|
|---|
scenario 2
Agent crash (or force killing the agent process) | - pid=$(ps -ef |grep cloudstack-agent |grep -v grep |awk '{print $2}') && kill -9 $pid
|
scenario 3 |
|---|
Agent is restarted manually
- systemctl restart cloudstack-agent
|
|---|
scenarios 5 (Needed ?)
mgmt server is restarted | - systemctl restart cloudstack-management
|
|---|
Migrate VM - MigrateVMCommand
- MigrateCommand (internal)
| intermittent failure - agent updates answer when connection is back to normal (every ping.interval in PingRoutingCommand)
- mgmt server saves the answer into DB
- mgmt server gets the answer from DB every 10 seconds
- mgmt server processes the answer if found
- Operation times out
continuous failure - agent
- Migration thread of VM [i-2-1198-VM] finished.
- management
- Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance
- Behavior:
- Active migration command so scheduling a restart
- Current: vm is Stopped on destination and Running on source (actually it is not if migration is successful)
- New: vm is Running on destination host if found (via libvirt)
- (new) check via destination host (if dest is Up)
- if vm is Running on destination, consider the vm migration is successful.
| - agent
- agent is started again by systemctl
- Note:
- In edge case, if the vm is migrated before agent is restarted (in 10 seconds) , the vm is Stopped (by ACS on destination host)
- solved
| - agent
- abort the migration job by disconnect hook (via libvirt)
- written by Marcus
- triggered when agent is restarted or disconnected
- tested with multiple mgmt servers
- management server
- Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance
- (new) check via destination host (if dest is Up)
- if vm is Running on destination, consider the vm migration is successful.
- Current: VM is set to Stopped state on destination host and then Running on the source host
- New:
- vm is Running on destination if found
- if not found, set to Stopped state on destination host and then Running on the source host (same as current)
- Note:
- In edge case, the vm is migrated before aborting the job, the vm is Stopped (by ACS on destination host)
- solved
| - mgmt server
- Cancel left-over job-17783
- Cleaning up Instance with Id: 1226
- CleanUp Async Jobs after mgmt server maintenance (#8394)
- vm is Running on source host
- agent (new)
- abort the migration job by disconnect hook
- written by Marcus
- Edge case: both scenario 1 and 5 happen
- No issues found, ALL looks ok
|
Migrate VM with volumes - (between NFS)
- MigrateVirtualMachine
WithVolumeCmd - MigrateCommand (internal)
Note: this action is not supported on PowerFlex | intermittent failure
continuous failure - current: vm is Stopped on destination and Running on source and pool (actually it is not if migration is successful)
- New: vm is Running on destination host and pool if found on destination host
| - management server
- Failed to migrate VM [VM instance xxx along with its volumes due to [com.cloud.utils.exception.CloudRuntimeException: Copy volume(s) to storage(s) xxx failed in StorageSystemDataMotionStrategy.copyAsync. Error message: [Commands 4446460207098233056 to Host 1 timed out after 21600].].
- Current:
- vm is Running on source host
- volume is Ready on source pool
- new volume is Migrating on destination pool (solved)
- new
- check via destination host (if dest is Up)
- if vm is Running on destination, consider the vm migration is successful.
- (after fix) new volume is Destroy on destination pool if fail
- Note:
- In edge case, if the vm is migrated before agent is restarted (in 10 seconds) , the vm is Stopped (by ACS on destination host)
- solved
| same as left | - management server
- vm is Running on source host
- both volumes are stuck at Migrating state (need to reconcile)
- Edge case: both scenario 1 and 5 happen
- VM is migrated but state is Running on source
How to reconcile (Verified): - If VM is Running state and volume is Migrating state
- check via host (if source is Up)
- determine the state and update
- if Running on dest, Running
- If Paused on dest, Migrating
- check disk path and update volume state
- if source volume is found but destination is not found, mark source as Ready and dest as Destroy
- if source volume is not found but destination is found, mark source as Destroy and dest as Ready
|
Migrate Volume (on NFS) - CopyCommand (from primary1 to sec1)
- CopyCommand (from sec1 to primary2)
ROOT/DATA volume of Stopped VM | intermittent failure - 1st CopyCommand: wait until connection is recoved
- 2nd CopyCommand: wait until connection is recoved
continuous failure - 1st CopyCommand: operation timeout
- Resource [StoragePool:1] is unreachable: Volume [{"name":"ROOT-1159","uuid":"ed322109-8e54-47ca-8183-3c8f2a0f37e6"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 5302988561228766380 to Host 1 timed out after 1800].
- new volume is Creating with empty path on destination pool (solved)
- 2nd CopyCommand
- Resource [StoragePool:1] is unreachable: Volume [{"name":"ROOT-1159","uuid":"ed322109-8e54-47ca-8183-3c8f2a0f37e6"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 7457960982925541434 to Host 1 timed out after 2400].
- new volume is Creating with empty path on destination pool (solved)
New - new volume is Destroy state with empty path on destination pool
| - 1st CopyCommand
- copy object failed: com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 15481123719087037 to Host 1 timed out after 1800
- new volume is Destroy state
- expungeVolumeAsync is called in copyVolumeCallBack
- 2nd CopyCommand
- copy failed com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 398850041998999773 to Host 1 timed out after 2400
- new volume is Destroy state
- expungeVolumeAsync is called in copyVolumeCallBack
- No issues found, ALL looks good
| same as left | - 1st CopyCommand
- new volume is Creating state
- record in volume_store_ref is Creating state (to be cleaned)
- 2nd CopyCommand
- new volume is Creating state
- record in volume_store_ref is Copying state (to be cleaned)
How to reconcile (verified): - 1st CopyCommand
- new volume is Destroy state and removed=now()
- volume_store_ref with Creating state is removed
- 2nd CopyCommand
- check the size of new volume. If size is changed, the process is still ongoing.
- if size is not changed, new volume is Destroy state and path is set (but not removed)
- volume_store_ref with Copying state is removed
|
Migrate Volume of Running VM (on Powerflex) |
TODO |
TODO |
TODO |
TODO |
Migrate Volume of Stopped VM (on Powerflex) |
TODO |
TODO |
TODO |
TODO |