...
| Code Block |
|---|
|
-- Add table for reconcile commands
CREATE TABLE IF NOT EXISTS `cloud`.`reconcile_commands` (
`id` bigint unsigned NOT NULL UNIQUE AUTO_INCREMENT,
`management_server_id` bigint unsigned NOT NULL COMMENT 'node id of the management server',
`host_id` bigint unsigned NOT NULL COMMENT 'id of the host',
`request_sequence` bigint unsigned NOT NULL COMMENT 'sequence of the request',
`state_by_management` varchar(255) COMMENT 'state of the command updated by management server',
`state_by_agent` varchar(255) COMMENT 'state of the command updated by cloudstack agent',
`command_name` varchar(255) COMMENT 'name of the command',
`command_info` MEDIUMTEXT COMMENT 'info of the command',
`answer_name` varchar(255) COMMENT 'name of the answer',
`answer_info` MEDIUMTEXT COMMENT 'info of the answer',
`created` datetime COMMENT 'date the reconcile command was created',
`removed` datetime COMMENT 'date the reconcile command was removed',
`updated` datetime COMMENT 'date the reconcile command was updated',
`retry_count` bigint unsigned DEFAULT 0 COMMENT 'The retry count of reconciliation',
PRIMARY KEY(`id`),
INDEX `i_reconcile_command__host_id`(`host_id`),
CONSTRAINT `fk_reconcile_command__host_id` FOREIGN KEY (`host_id`) REFERENCES `host`(`id`) ON DELETE CASCADE
) ENGINE=InnoDB DEFAULT CHARSET=utf8;
-- Add last_id to the volumes table
CALL `cloud`.`IDEMPOTENT_ADD_COLUMN`('cloud.volumes', 'last_id', 'bigint(20) unsigned DEFAULT NULL'); |
3.3 Core changes
Command.java
| Code Block |
|---|
|
public enum State {
CREATED, // Command is created by management server
STARTED, // Command is started by agent
PROCESSING, // Processing by agent
PROCESSING_IN_BACKEND, // Processing in backend by agent
COMPLETED, // Operation succeeds by agent or management server
FAILED, // Operation fails by agent
RECONCILE_READYRETRY, // Ready for reconciliation
RECONCILING, // Being reconciled by management server
RECONCILED, // Reconciled by management server
RECONCILE_FAILED, // Fail to reconcile by management server
TIMED_OUT, // Timed out on management server or agent
INTERRUPTED, // Interrupted by management server or agent (for example agent is restarted),
DANGLED_IN_BACKEND // Backend process which cannot be processed normally (for example agent is restarted)
} |
...
3.4 Considerations
| Scenarios | Current behavior | How to deal with it |
|---|
- Scenario 1: Connection failures between the agent and the management server
| Agent cannot send Answer to management server, therefore timed out on management server | Agent: (for reconcile command only) - Save command info as JSON file
- Update command state in JSON when start/process/complete the command
- STARTED
- PROCESSING
- COMPLETE/FAILED
- Save Answer in the JSON file
- Every minute
- Load command/answer from JSON files
- Send with with PingCommand
- Receive the PingAnswer from management server
- Remove the JSON file if state is COMPLETE/FAILED
Management server: (for reconcile command only) - When create the reconcile command, insert a record with state CREATED into reconcile_commands table
- When receive the PingCommand, Update command state and answer in reconcile_commands table
- When wait for the answer of command (reconcile command only)
- Every 10 seconds, check answer of the reconcile_commands table
- If answer is found, parse the answer
- Returna and continue with the answer
- No need to wait until timeout, if there are connection failures.
- if timed out, update command state to TIMED_OUT in reconcile_commands table
|
- Scenario 2: Agent crash (or force killing the agent process)
| Agent interrupt the process or processing in backend | Management server - When Host Status is determined as Down
- Update reconcile commands (to the agent) in reconcile_command table
- to INTERRUPTED state
- Reconcile the command
- see "3.6 Reconcile the command"
|
- Scenario 3: Agent restart
| Agent interrupt the process or processing in backend | Agent (when restart) When stop the agent, updates state of processesPROCESSING to INTERRUPTEDPROCESSING_IN_BACKEND to DANGLED_IN_BACKEND
- When start the agent, updates state of processes
- PROCESSING to INTERRUPTED
- PROCESSING_IN_BACKEND to DANGLED_IN_BACKEND
- Agent send CommandInfo to management server via PingCommand every minute
- with new state
- Management server reconcile commands every minute (from reconcile_commands table) , for commands in state
|
- Scenario 4: Agent has completed but timed out
| timed out on management server | Management server - Update state_by_management to TIMED_OUT
- Update state_by_agent to COMPLETED (from PROCESSING)
- Reconcile the command
- see "3.6 Reconcile the command"
- TODO: what should be the correct resource state ?4
|
- Scenario 5: Management restart
| Agent process the command - has completed and send Answer to management server, but no action on management server
- has not completed, and is processing command
| Management server - Update state_by_management to INTERRUPTED when mgmt server is stopped.
- When another management server is detected DOWN
- update reconcile_command to INTERRUPTED state by management server
|
3.5 State transitions (Agent)
Management server (all good): CREATED -> COMPLETED
...
States: INTERRUPTED/TIMED_OUT/RECONCILE_READYRETRY/RECONCILE_FAILED
-> RECONCILING
-> RECONCILED (all good) / RECONCILE_FAILED (failed, will retry) / RECONCILED_READY RETRY (success, but need more information, will retry)
Migrate VM without volumes (If VM is Migrating)
- PrepareForMigrationCommand (Update VM state)
- MigrateCommand
- check via source host (if source is Up)
- check via destination host (if dest is Up)
- determine the state and update
Migrate VM with volumes (If VM is Migrating)
- check VM state and volume states
- Check if there are reconcile commands for the VM
- If no, check the state on source host (last_host_id) and destination host (host_id)
- PrepareForMigrationCommand (Update VM state)
- MigrateCommand
- check via source host (if source is Up)
- check via destination host (if dest is Up)
- determine the state and update
Migrate Volumes (If volume is Migrating)
- CopyCommand (from primary1 to secondary)
- skipped.
- check if there are other Command on same volume ?
- CopyCommand (from secondary to primary2)
- check if volume exists on primary2 (via the host, other host on same cluster or pod or zone, depends on the scope of storage, or cluster of host) TODO
- check if volume is changed on primary2
- if yes, still Copying
- if no, update state
- CopyCommand (from primary1 to primary2)
- check if volume exists on primary1
- check if volume exists on primary2
- check if volume is changed on primary2
- if yes, still Copying
- if no, update state
- MigrateVolumeCommand (from primary1 to primary2)
- check if volume exists on primary1
- check if volume exists on primary2
- check if volume is changed on primary2
- if yes, still Migrating
- if no, update state
3.7 Limitations
- Assumption
- Each request does not have multiple commands with same name
- Only support:
- CopyCommand, MigrateCommand, MigrateVolumeCommand
- resources in Migration state (vm, source volume, dest volume)
- Hypervisor: KVM
- Storage: NFS, Local, Powerflex
3.8 To be Discussed (TODO)
- How to distribute the reconciiation tasks if there are multiple management servers ?
- How to better handle the state DANGLED_IN_BACKEND ?
- How to better handle the state COMPLETED and FAILED ? Do not reconcile via hosts if state_by_agent to COMPLETED ?
4. Test cases
Please note:
Migration between NFS and Local requires the fix: https://github.com/apache/cloudstack/pull/10266
See "4.3 Summary of test results" on how to reconcile the command
3.7 Limitations
- Assumption
- Each request does not have multiple commands with same name
- for example, 700028267079401857-org.apache.cloudstack.storage.command.CopyCommand
- Only support:
- 3 commands
- CopyCommand
- MigrateCommand
- MigrateVolumeCommand
- resources in Migration state (vm, source volume, dest volume)
- skipped if the resource state is not Migrating
- skipped if the source or dest host/pool are inconsistent with command
- Hypervisor: KVM
- Storage: NFS, Local, Powerflex
3.8 To be Discussed (TODO)
- How to distribute the reconciiation tasks if there are multiple management servers ?
- now: Reconciliation is processed by first management server
- How to better handle the state DANGLED_IN_BACKEND ?
- now: no special action
- How to better handle the state COMPLETED and FAILED ? Do not reconcile via hosts if state_by_agent to COMPLETED ?
- now: no special action
4. Test cases
4.1 Backend commands of VM and volume migrations
Please note:
Migration between NFS and Local requires the fix: https://github.com/apache/cloudstack/pull/10266
| NFS to NFS | NFS to Local | Local to Local | Local to NFS | Powerflex to Powerflex (different envs) | Powerflex <------> NFS |
|---|
Migrate VM (NFS or Powerflex) | PrepareForMigrationCommand (dest) MigrateCommand (source) | - | - | - | PrepareForMigrationCommand (dest) MigrateCommand (source) | - |
Migrate VM with volumes
(NFS only) | CopyCommand (template to primary if needed) CreateObjectCommand (new volume) ModifyTargetsCommand PrepareForMigrationCommand (dest) MigrateCommand (source) DeleteCommand (source) | same as "NFS to NFS" | same as "NFS to NFS" | same as "NFS to NFS" | Migrating a volume online with KVM from managed storage is not currently supported. | Pool [%s] is not compatible with volume [%s], skipping it. |
Migrate ROOT Volume (of Running VM)
(Powerflex only) | KVM does not support volume live migrationdue to the limited possibility to refresh VM XML domain. Therefore, to live migrate a volume between storage pools, one must migrate the VM to a different host as well to force the VM XML domain update. Use 'migrateVirtualMachineWithVolumes' instead. | same | same | same | MigrateVolumeCommand | Storage pool pr503-t11980-kvm-ol8-kvm-pri3 is not suitable to migrate volume |
Migrate ROOT Volume (of Stopped VM)
(NFS or Powerflex) | CopyCommand (primary1 to secondary) CopyCommand (secondary to primary2) DeleteCommand (secondary) DeleteCommand (primary1) | same | same | same | CopyCommand (primary1 to primary2) | same as above |
|
|
|
|
|
|
|
| Migrate DATA |
NFS to NFS | NFS to Local | Local to Local | Local to NFS | Powerflex to Powerflex | Powerflex <------> NFS | | Migrate VM | PrepareForMigrationCommand (dest) MigrateCommand (source) | - | - | - | PrepareForMigrationCommand (dest) MigrateCommand (source) | - |
| Migrate VM with volumes | CopyCommand (template to primary if needed) CreateObjectCommand (new volume) ModifyTargetsCommand PrepareForMigrationCommand (dest) MigrateCommand (source) DeleteCommand (source) | same as "NFS to NFS" | same as "NFS to NFS" | same as "NFS to NFS" | Migrating a volume online with KVM from managed storage is not currently supported. | Pool [%s] is not compatible with volume [%s], skipping it. |
| Migrate ROOT Volume (of Running VM) | KVM does not support volume live migrationdue to the limited possibility to refresh VM XML domain. Therefore, to live migrate a volume between storage pools, one must migrate the VM to a different host as well to force the VM XML domain update. Use 'migrateVirtualMachineWithVolumes' instead. | same | same | same | MigrateVolumeCommand | Storage pool pr503-t11980-kvm-ol8-kvm-pri3 is not suitable to migrate volume | | same as above |
| Migrate DATA Migrate ROOT Volume (of Stopped VM) | CopyCommand (primary1 to secondary) CopyCommand (secondary to primary2) DeleteCommand (secondary) DeleteCommand (primary1) | same | same | same | CopyCommand (primary1 to primary2) | same as above |
| Migrate DATA Volume (of Running VM)KVM does not support volume live migrationdue to the limited possibility to refresh VM XML domain. Therefore, to live migrate a volume between storage pools, one must migrate the VM to a different host as well to force the VM XML domain update. Use 'migrateVirtualMachineWithVolumes' instead.unattached) | CopyCommand (primary1 to secondary) CopyCommand (secondary to primary2) DeleteCommand (secondary) DeleteCommand (primary1) | same | same | same | CopyCommand (primary1 to primary2) MigrateVolumeCommand | same as above |
| Migrate DATA Volume (of Stopped VM) | CopyCommand (primary1 to secondary) CopyCommand (secondary to primary2) DeleteCommand (secondary) DeleteCommand (primary1) | same | same | same | CopyCommand (primary1 to primary2) | same as above |
| Migrate DATA Volume (unattached) | CopyCommand (primary1 to secondary) CopyCommand (secondary to primary2) DeleteCommand (secondary) DeleteCommand (primary1) | same | same | same | CopyCommand (primary1 to primary2) | same as above |
4.2 How to test
...
- systemctl restart cloudstack-agent
...
- systemctl restart cloudstack-management
4.2 How to test
- agent is restarted
- systemctl restart cloudstack-agent
- management server is restarted
- systemctl restart cloudstack-management
- agent and management server communication failure
- iptables -I OUTPUT -p tcp -m tcp --dport 8250 -j DROP
- iptables -D OUTPUT -p tcp -m tcp --dport 8250 -j DROP
- Agent crash
- pid=$(ps -ef |grep cloudstack-agent |grep -v grep |awk '{print $2}')
- kill -9 $pid
- Agent has completed but timed out
4.3 Summary of test results
| Action | scenario 1: Connection failures between agent and management server Scenario 4: Agent has completed but timed out | scenario 2 Agent crash (or force killing the agent process) | scenario 3 Agent is restarted manually | scenarios 5 (Needed ?) mgmt server is restarted |
|---|
| How to test |
|---|
...
| - iptables -I OUTPUT -p tcp -m
|
|---|
...
- iptables -D OUTPUT -p tcp
|
|---|
...
- -m tcp --dport 8250 -j DROP
|
|---|
...
- pid=$(ps -ef |grep cloudstack-agent |grep -v grep |awk '{print $2}')
|
|---|
...
...
| - systemctl restart cloudstack-agent
| - systemctl restart cloudstack-management
|
|---|
1. Migrate VM - MigrateVMCommand
- MigrateCommand (internal)
Supported by NFS or Powerflex | intermittent failure - agent updates answer when connection is back to normal (every reconcile.command.interval in PingRoutingCommand)
- mgmt server saves the answer into DB
- mgmt server gets the answer from DB every 10 seconds
- mgmt server processes the answer if found
- Operation times out
continuous failure - agent
- Migration thread of VM [i-2-1198-VM] finished.
- management
- Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance
- Behavior:
- Active migration command so scheduling a restart
- Current: vm is Stopped on destination and Running on source (actually it is not if migration is successful)
- New: vm is Running on destination host if found (via libvirt)
- (new) check via destination host (if dest is Up)
- if vm is Running on destination, consider the vm migration is successful.
| - agent
- agent is started again by systemctl
- Note:
- In edge case, if the vm is migrated before agent is restarted (in 10 seconds) , the vm is Stopped (by ACS on destination host)
- solved
| - agent
- when agent is restarted or disconnected
- abort the migration job by a disconnect hook (MigrationCancelHook)
- written by Marcus
- tested with multiple mgmt servers
- management server
- Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance
- (new) check via destination host (if dest is Up)
- if vm is Running on destination, consider the vm migration is successful.
- Current: VM is set to Stopped state on destination host and then Running on the source host
- New:
- vm is Running on destination if found
- if not found, set to Stopped state on destination host and then Running on the source host (same as current)
- Note:
- In edge case, the vm is migrated before aborting the job, the vm is Stopped (by ACS on destination host)
- solved
| Issues - mgmt server
- Cancel left-over job-17783
- Cleaning up Instance with Id: 1226
- CleanUp Async Jobs after mgmt server maintenance (#8394)
- vm is Running on source host (this will be updated when mgmt server gets vm report from agent)
- agent (new)
- abort the migration job by disconnect hook
- written by Marcus
- Edge case: both scenario 1 and 5 happen
- No issues found, ALL looks ok
|
2. Migrate VM with volumes - (between NFS)
- MigrateVirtualMachine
WithVolumeCmd - MigrateCommand (internal)
Note: this action is not supported on PowerFlex
Two volumes in DB - source volume
- dest volume (last_id = source_volume_id)
| intermittent failure
continuous failure - current: vm is Stopped on destination and Running on source and pool (actually it is not if migration is successful)
- New: vm is Running on destination host and pool if found on destination host
| - management server
- Failed to migrate VM [VM instance xxx along with its volumes due to [com.cloud.utils.exception.CloudRuntimeException: Copy volume(s) to storage(s) xxx failed in StorageSystemDataMotionStrategy.copyAsync. Error message: [Commands 4446460207098233056 to Host 1 timed out after 21600].].
- Current:
- vm is Running on source host
- volume is Ready on source pool
- Bug: new volume is Migrating on destination pool (solved)
- new
- check via destination host (if dest is Up)
- if vm is Running on destination, consider the vm migration is successful.
- (after fix) new volume is Destroy on destination pool if fail
- Note:
- In edge case, if the vm is migrated before agent is restarted (in 10 seconds) , the vm is Stopped (by ACS on destination host)
- solved
| same as left | Issues - management server
- vm is Running on source host
- Bug: both volumes are stuck at Migrating state (need to reconcile)
- Edge case: both scenario 1 and 5 happen
- VM is migrated but state is Running on source
How to reconcile (Verified): - only if
- VM is Running state
- volume is Migrating state
- check via host (if source is Up)
- check disk path and update volume state
- if source volume is found but destination is not found, mark source as Ready and dest as Destroy
- if source volume is not found but destination is found, mark source as Destroy and dest as Ready
|
3. Migrate Volume (on NFS) - CopyCommand (from primary1 to sec1)
- CopyCommand (from sec1 to primary2)
ROOT/DATA volume of Stopped VM | intermittent failure - 1st CopyCommand: wait until connection is recoved
- 2nd CopyCommand: wait until connection is recoved
continuous failure - 1st CopyCommand: operation timeout
- Resource [StoragePool:1] is unreachable: Volume [{"name":"ROOT-1159","uuid":"ed322109-8e54-47ca-8183-3c8f2a0f37e6"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 5302988561228766380 to Host 1 timed out after 1800].
- Bug: new volume is Creating with empty path on destination pool (solved)
- 2nd CopyCommand
- Resource [StoragePool:1] is unreachable: Volume [{"name":"ROOT-1159","uuid":"ed322109-8e54-47ca-8183-3c8f2a0f37e6"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 7457960982925541434 to Host 1 timed out after 2400].
- Bug: new volume is Creating with empty path on destination pool (solved)
New - new volume is Destroy state with empty path on destination pool
| - 1st CopyCommand
- copy object failed: com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 15481123719087037 to Host 1 timed out after 1800
- new volume is Destroy state
- expungeVolumeAsync is called in copyVolumeCallBack
- 2nd CopyCommand
- copy failed com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 398850041998999773 to Host 1 timed out after 2400
- new volume is Destroy state
- expungeVolumeAsync is called in copyVolumeCallBack
- No issues found, ALL looks good
| same as left | Issues - 1st CopyCommand
- new volume is Creating state
- record in volume_store_ref is Creating state (to be cleaned)
- 2nd CopyCommand
- new volume is Creating state
- record in volume_store_ref is Copying state (to be cleaned)
How to reconcile (verified): - 1st CopyCommand
- new volume is Destroy state and removed=now()
- volume_store_ref with Creating state is removed
- 2nd CopyCommand
- check the size of new volume. If size is changed, the process is still ongoing.
- if size is not changed, new volume is Destroy state and path is set (but not removed)
- volume_store_ref with Copying state is removed
|
4. Migrate Volume of Running VM (on Powerflex)
| - agent: succeed
- answer: {"volumePath":"1fe84f0700000002:vol-610-4d48-306f","result":true,"contextMap":{},"wait":0,"bypassHostMaintenance":false}
- management server
- volume path is set to new, then set to old when fail
- Bug: Actually, on KVM host, the VM is running with new volume, but the new volume is removed
- Input/Output error, read-only system
- New:
- check volume statistics on destination pool via ScaleIO gateway
- if allocation size = provisioned size, volume migration is done. remove volume on source pool.
- if allocation size is not same to provisioned size, volume migration failes, remove volume on destination pool.
| - management
- Resource [StoragePool:12] is unreachable: Migrate volume failed: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 2520889891420635168 to Host 1 timed out after 2400
- Failed to migrate PowerFlex volume
- volume path is set to new, then set to old when fail (looks good)
- revertBlockCopyVolumeOperations
- No issues found, ALL looks good
changes on agent - it uses "dm.blockCopy" (introduced by Hari)
- when agent is disconnected or restarted
- abort the block copy job by a disconnect hook (VolumeMigrationCancelHook)
- similar as MigrationCancelHook
| same as left | Issue: - one volume in Migrating in DB
- path, iscsi_name, pool_id are info on destination pool
- when mgmt server is restarted, it updates all Migrating volume to Ready
- this is disabled if reconcile is enabled
How to reconcile: - only if vm is Running and volume in Migrating state
- The current pool is destination pool of reconcile command
- check vm state and vm disks via agent
- if volume is Ready on source pool,
- revert path, iscsi_name and pool_id to source pool
- create a dummy volume on dest pool
- remove the dummy volume
- if volume is not found on source pool
- update state to Ready
- create a dummy volume on source pool
- remove the dummy volume
- if vm is running on source pool (check by fullpath) = volume is ready on source pool
- if vm is running on dest pool, update volume state to Ready, and remove volume from source pool
|
5. Migrate Volume of Stopped VM (on Powerflex)
| intermittent failure - CopyCommand: wait until connection is recoved
continuous failure - Resource [StoragePool:11] is unreachable: Volume [{"name":"ROOT-533","uuid":"ff8a53a0-d513-4bfc-8920-352948ea066b"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 8886446489732120778 to Host 1 timed out after 2400]
- 2400 = wait (1200) * 2
- new volume is Expunged
- due to expungeVolumeAsync in copyManagedVolumeCallBack
- Not an issue since the VM is Stopped so no data loss
| - Resource [StoragePool:12] is unreachable: Volume [{"name":"ROOT-531","uuid":"4225516d-52fe-465d-a46d-81f97a2e55b2"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 3050062847636668430 to Host 1 timed out after 2400].
- new volume is Expunged
- due to expungeVolumeAsync in copyManagedVolumeCallBack
- No issues found, ALL looks good
| same as left | Issue: two volumes in DB - source volume: Migrating state
- dest volume: Creating state
- dest.last_id = source_volume_id
How to reconcile - check volume information via host
- determine state of source and dest volumes
- if source volume is Ready
- update source from Migrating to Ready
- update dest volume from Creating to Destroy
- if dest volume is Ready
- update dest from Creating to Ready
- update source volume from Migrating to Destroy
TO be fixed - when migrate volume from powerflex to powerflex with same system id, it will not create CopyCommand
|