...
| Code Block |
|---|
|
-- Add table for reconcile commands
CREATE TABLE IF NOT EXISTS `cloud`.`reconcile_commands` (
`id` bigint unsigned NOT NULL UNIQUE AUTO_INCREMENT,
`management_server_id` bigint unsigned NOT NULL COMMENT 'node id of the management server',
`host_id` bigint unsigned NOT NULL COMMENT 'id of the host',
`request_sequence` bigint unsigned NOT NULL COMMENT 'sequence of the request',
`state_by_management` varchar(255) COMMENT 'state of the command updated by management server',
`state_by_agent` varchar(255) COMMENT 'state of the command updated by cloudstack agent',
`command_name` varchar(255) COMMENT 'name of the command',
`command_info` MEDIUMTEXT COMMENT 'info of the command',
`answer_name` varchar(255) COMMENT 'name of the answer',
`answer_info` MEDIUMTEXT COMMENT 'info of the answer',
`created` datetime COMMENT 'date the reconcile command was created',
`removed` datetime COMMENT 'date the reconcile command was removed',
`updated` datetime COMMENT 'date the reconcile command was updated',
`retry_count` bigint unsigned DEFAULT 0 COMMENT 'The retry count of reconciliation',
PRIMARY KEY(`id`),
INDEX `i_reconcile_command__host_id`(`host_id`),
CONSTRAINT `fk_reconcile_command__host_id` FOREIGN KEY (`host_id`) REFERENCES `host`(`id`) ON DELETE CASCADE
) ENGINE=InnoDB DEFAULT CHARSET=utf8;
-- Add last_id to the volumes table
CALL `cloud`.`IDEMPOTENT_ADD_COLUMN`('cloud.volumes', 'last_id', 'bigint(20) unsigned DEFAULT NULL'); |
3.3 Core changes
Command.java
| Code Block |
|---|
|
public enum State {
CREATED, // Command is created by management server
STARTED, // Command is started by agent
PROCESSING, // Processing by agent
PROCESSING_IN_BACKEND, // Processing in backend by agent
COMPLETED, // Operation succeeds by agent or management server
FAILED, // Operation fails by agent
RECONCILE_READYRETRY, // Ready for reconciliation
RECONCILING, // Being reconciled by management server
RECONCILED, // Reconciled by management server
RECONCILE_FAILED, // Fail to reconcile by management server
TIMED_OUT, // Timed out on management server or agent
INTERRUPTED, // Interrupted by management server or agent (for example agent is restarted),
DANGLED_IN_BACKEND // Backend process which cannot be processed normally (for example agent is restarted)
} |
new method for each Command
| Code Block |
|---|
|
public boolean isReconcile() {
return false;
} |
...
3.4 Considerations
| Scenarios | Current behavior | How to deal with it |
|---|
- Scenario 1: Connection failures between the agent and the management server
| Agent cannot send Answer to management server, therefore timed out on management server | Agent: (for reconcile command only) - Save command info as JSON file
- Update command state in JSON when start/process/complete the command
- STARTED
- PROCESSING
- COMPLETE/FAILED
- Save Answer in the JSON file
- Every minute
- Load command/answer from JSON files
- Send with with PingCommand
- Receive the PingAnswer from management server
- Remove the JSON file if state is COMPLETE/FAILED
Management server: (for reconcile command only) - When create the reconcile command, insert a record with state CREATED into reconcile_commands table
- When receive the PingCommand, Update command state and answer in reconcile_commands table
- When wait for the answer of command (reconcile command only)
- Every 10 seconds, check answer of the reconcile_commands table
- If answer is found, parse the answer
- Returna and continue with the answer
- No need to wait until timeout, if there are connection failures.
- if timed out, update command state to TIMED_OUT in reconcile_commands table
|
- Scenario 2: Agent crash (or force killing the agent process)
| Agent interrupt the process or processing in backend | Management server |
(TODO)- When Host Status is determined as Down
|
(only cloudstack-agent is DOWN?)- Update reconcile commands (to the agent) in reconcile_command table
- to INTERRUPTED state
- Reconcile the command
- see "3.6 Reconcile the command"
|
Update state to RECONCILED, if command is reconciled and no need more reconciliation- Scenario 3: Agent restart
| Agent interrupt the process or processing in backend |
) | Agent (when restart) When stop the agent, updates state of
|
processes (TODO)processesPROCESSING to INTERRUPTEDPROCESSING_IN_BACKEND to DANGLED_IN_BACKEND
- When start the agent, updates state of processes
- PROCESSING to INTERRUPTED
- PROCESSING_IN_BACKEND to DANGLED_IN_BACKEND
- Agent send CommandInfo to management server via PingCommand every minute
- with new state
- Management server reconcile commands every minute (from reconcile_
|
command - commands table) , for commands in state
|
- Scenario 4: Agent has completed but timed out
| timed out on management server | Management server - Update state_by_management to TIMED_OUT
- Update state_by_agent to COMPLETED (from PROCESSING)
- Reconcile the command
- see "3.6 Reconcile the command"
- TODO: what should be the correct resource state ?4
|
- Scenario 5: Management restart
| Agent process the command - has completed and send Answer to management server, but no action on management server
- has not completed, and
|
process | Management server - Update state_by_management to INTERRUPTED when mgmt server is stopped.
- When another management server is detected DOWN
|
, - update reconcile_command to INTERRUPTED state
|
? (TODO)
3.5 State transitions (Agent)
Management server (all good): CREATED -> COMPLETED
...
Agent (all good): STARTED → PROCESSING → COMPLETED → Send Answer to management server → Mgmt server removes updates the reconcile command jobs (new table: reconcile_command ?) → commands ) , and removes if COMPLETED→ Agent removes the JSON file if job is not found (TODO).or COMPLETED or FAILED
Agent (wrong): STARTED → PROCESSING → FAILED → Send Answer to management server → Mgmt server removes the reconcile command jobs (new table: reconcile_command ?commands) → Agent removes the JSON file if job is not found.
...
States: INTERRUPTED/TIMED_OUT/RECONCILE_READYRETRY/RECONCILE_FAILED
-> RECONCILING
-> RECONCILED (all good) / RECONCILE_FAILED (failed, will retry) / RECONCILED_READY RETRY (success, but need more information, will retry)
- Migrate VM: via destination host (TODO)
- Migrate Volume: via other host (TODO)
- Copy Command: via other host on same cluster or pod or zone, depends on the scope of storage, or cluster of host (TODO)
4. Test cases
See "4.3 Summary of test results" on how to reconcile the command
3.7 Limitations
- Assumption
- Each request does not have multiple commands with same name
- for example, 700028267079401857-org.apache.cloudstack.storage.command.CopyCommand
- Only support:
- 3 commands
- CopyCommand
- MigrateCommand
- MigrateVolumeCommand
- resources in Migration state (vm, source volume, dest volume)
- skipped if the resource state is not Migrating
- skipped if the source or dest host/pool are inconsistent with command
- Hypervisor: KVM
- Storage: NFS, Local, Powerflex
3.8 To be Discussed (TODO)
- How to distribute the reconciiation tasks if there are multiple management servers ?
- now: Reconciliation is processed by first management server
- How to better handle the state DANGLED_IN_BACKEND ?
- now: no special action
- How to better handle the state COMPLETED and FAILED ? Do not reconcile via hosts if state_by_agent to COMPLETED ?
- now: no special action
4. Test cases
4.1 Backend commands of VM and volume migrations
Please note:
Migration between NFS and Local requires the fix: https://github.com/apache/cloudstack/pull/10266
| NFS to NFS | NFS to Local | Local to Local | Local to NFS | Powerflex to Powerflex (different envs) | Powerflex <------> NFS |
|---|
Migrate VM (NFS or Powerflex) | PrepareForMigrationCommand (dest) MigrateCommand (source) | - | - | - | PrepareForMigrationCommand (dest) MigrateCommand (source) | - |
Migrate VM with volumes
(NFS only) | CopyCommand (template to primary if needed) CreateObjectCommand (new volume) ModifyTargetsCommand PrepareForMigrationCommand (dest) MigrateCommand (source) DeleteCommand (source) | same as "NFS to NFS" | same as "NFS to NFS" | same as "NFS to NFS" | Migrating a volume online with KVM from managed storage is not currently supported. | Pool [%s] is not compatible with volume [%s], skipping it. |
Migrate ROOT Volume (of Running VM)
(Powerflex only) | KVM does not support volume live migrationdue to the limited possibility to refresh VM XML domain. Therefore, to live migrate a volume between storage pools, one must migrate the VM to a different host as well to force the VM XML domain update. Use 'migrateVirtualMachineWithVolumes' instead. | same | same | same | MigrateVolumeCommand | Storage pool pr503-t11980-kvm-ol8-kvm-pri3 is not suitable to migrate volume |
Migrate ROOT Volume (of Stopped VM)
(NFS or Powerflex) | CopyCommand (primary1 to secondary) CopyCommand (secondary to primary2) DeleteCommand (secondary) DeleteCommand (primary1) | same | same | same | CopyCommand (primary1 to primary2) | same as above |
|
|
|
|
|
|
|
| Migrate DATA Volume (of Running VM) | KVM does not support volume live migrationdue to the limited possibility to refresh VM XML domain. Therefore, to live migrate a volume between storage pools, one must migrate the VM to a different host as well to force the VM XML domain update. Use 'migrateVirtualMachineWithVolumes' instead. | same | same | same | MigrateVolumeCommand | same as above |
| Migrate DATA Volume (of Stopped VM) | CopyCommand (primary1 to secondary) CopyCommand (secondary to primary2) DeleteCommand (secondary) DeleteCommand (primary1) | same | same | same | CopyCommand (primary1 to primary2) | same as above |
| Migrate DATA Volume (unattached) | CopyCommand (primary1 to secondary) CopyCommand (secondary to primary2) DeleteCommand (secondary) DeleteCommand (primary1) | same | same | same | CopyCommand (primary1 to primary2) | same as above |
4.2 How to test
- agent is restarted
- systemctl restart cloudstack-agent
- management server is restarted
- systemctl restart cloudstack-management
- agent and management server communication failure
- iptables -I OUTPUT -p tcp -m tcp --dport 8250 -j DROP
- iptables -D OUTPUT -p tcp -m tcp --dport 8250 -j DROP
- Agent crash
- pid=$(ps -ef |grep cloudstack-agent |grep -v grep |awk '{print $2}')
- kill -9 $pid
- Agent has completed but timed out
4.3 Summary of test results
| Action | scenario 1: Connection failures between agent and management server Scenario 4: Agent has completed but timed out | scenario 2 Agent crash (or force killing the agent process) | scenario 3 Agent is restarted manually | scenarios 5 (Needed ?) mgmt server is restarted |
|---|
| How to test | - iptables -I OUTPUT -p tcp -m tcp --dport 8250 -j DROP
- iptables -D OUTPUT -p tcp -m tcp --dport 8250 -j DROP
| - pid=$(ps -ef |grep cloudstack-agent |grep -v grep |awk '{print $2}') && kill -9 $pid
| - systemctl restart cloudstack-agent
| - systemctl restart cloudstack-management
|
|---|
1. Migrate VM - MigrateVMCommand
- MigrateCommand (internal)
Supported by NFS or Powerflex | intermittent failure - agent updates answer when connection is back to normal (every reconcile.command.interval in PingRoutingCommand)
- mgmt server saves the answer into DB
- mgmt server gets the answer from DB every 10 seconds
- mgmt server processes the answer if found
- Operation times out
continuous failure - agent
- Migration thread of VM [i-2-1198-VM] finished.
- management
- Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance
- Behavior:
- Active migration command so scheduling a restart
- Current: vm is Stopped on destination and Running on source (actually it is not if migration is successful)
- New: vm is Running on destination host if found (via libvirt)
- (new) check via destination host (if dest is Up)
- if vm is Running on destination, consider the vm migration is successful.
| - agent
- agent is started again by systemctl
- Note:
- In edge case, if the vm is migrated before agent is restarted (in 10 seconds) , the vm is Stopped (by ACS on destination host)
- solved
| - agent
- when agent is restarted or disconnected
- abort the migration job by a disconnect hook (MigrationCancelHook)
- written by Marcus
- tested with multiple mgmt servers
- management server
- Resource [Host:1] is unreachable: Host 1: Operation timed out on migrating VM instance
- (new) check via destination host (if dest is Up)
- if vm is Running on destination, consider the vm migration is successful.
- Current: VM is set to Stopped state on destination host and then Running on the source host
- New:
- vm is Running on destination if found
- if not found, set to Stopped state on destination host and then Running on the source host (same as current)
- Note:
- In edge case, the vm is migrated before aborting the job, the vm is Stopped (by ACS on destination host)
- solved
| Issues - mgmt server
- Cancel left-over job-17783
- Cleaning up Instance with Id: 1226
- CleanUp Async Jobs after mgmt server maintenance (#8394)
- vm is Running on source host (this will be updated when mgmt server gets vm report from agent)
- agent (new)
- abort the migration job by disconnect hook
- written by Marcus
- Edge case: both scenario 1 and 5 happen
- No issues found, ALL looks ok
|
2. Migrate VM with volumes - (between NFS)
- MigrateVirtualMachine
WithVolumeCmd - MigrateCommand (internal)
Note: this action is not supported on PowerFlex
Two volumes in DB - source volume
- dest volume (last_id = source_volume_id)
| intermittent failure
continuous failure - current: vm is Stopped on destination and Running on source and pool (actually it is not if migration is successful)
- New: vm is Running on destination host and pool if found on destination host
| - management server
- Failed to migrate VM [VM instance xxx along with its volumes due to [com.cloud.utils.exception.CloudRuntimeException: Copy volume(s) to storage(s) xxx failed in StorageSystemDataMotionStrategy.copyAsync. Error message: [Commands 4446460207098233056 to Host 1 timed out after 21600].].
- Current:
- vm is Running on source host
- volume is Ready on source pool
- Bug: new volume is Migrating on destination pool (solved)
- new
- check via destination host (if dest is Up)
- if vm is Running on destination, consider the vm migration is successful.
- (after fix) new volume is Destroy on destination pool if fail
- Note:
- In edge case, if the vm is migrated before agent is restarted (in 10 seconds) , the vm is Stopped (by ACS on destination host)
- solved
| same as left | Issues - management server
- vm is Running on source host
- Bug: both volumes are stuck at Migrating state (need to reconcile)
- Edge case: both scenario 1 and 5 happen
- VM is migrated but state is Running on source
How to reconcile (Verified): - only if
- VM is Running state
- volume is Migrating state
- check via host (if source is Up)
- check disk path and update volume state
- if source volume is found but destination is not found, mark source as Ready and dest as Destroy
- if source volume is not found but destination is found, mark source as Destroy and dest as Ready
|
3. Migrate Volume (on NFS) - CopyCommand (from primary1 to sec1)
- CopyCommand (from sec1 to primary2)
ROOT/DATA volume of Stopped VM | intermittent failure - 1st CopyCommand: wait until connection is recoved
- 2nd CopyCommand: wait until connection is recoved
continuous failure - 1st CopyCommand: operation timeout
- Resource [StoragePool:1] is unreachable: Volume [{"name":"ROOT-1159","uuid":"ed322109-8e54-47ca-8183-3c8f2a0f37e6"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 5302988561228766380 to Host 1 timed out after 1800].
- Bug: new volume is Creating with empty path on destination pool (solved)
- 2nd CopyCommand
- Resource [StoragePool:1] is unreachable: Volume [{"name":"ROOT-1159","uuid":"ed322109-8e54-47ca-8183-3c8f2a0f37e6"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 7457960982925541434 to Host 1 timed out after 2400].
- Bug: new volume is Creating with empty path on destination pool (solved)
New - new volume is Destroy state with empty path on destination pool
| - 1st CopyCommand
- copy object failed: com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 15481123719087037 to Host 1 timed out after 1800
- new volume is Destroy state
- expungeVolumeAsync is called in copyVolumeCallBack
- 2nd CopyCommand
- copy failed com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 398850041998999773 to Host 1 timed out after 2400
- new volume is Destroy state
- expungeVolumeAsync is called in copyVolumeCallBack
- No issues found, ALL looks good
| same as left | Issues - 1st CopyCommand
- new volume is Creating state
- record in volume_store_ref is Creating state (to be cleaned)
- 2nd CopyCommand
- new volume is Creating state
- record in volume_store_ref is Copying state (to be cleaned)
How to reconcile (verified): - 1st CopyCommand
- new volume is Destroy state and removed=now()
- volume_store_ref with Creating state is removed
- 2nd CopyCommand
- check the size of new volume. If size is changed, the process is still ongoing.
- if size is not changed, new volume is Destroy state and path is set (but not removed)
- volume_store_ref with Copying state is removed
|
4. Migrate Volume of Running VM (on Powerflex)
| - agent: succeed
- answer: {"volumePath":"1fe84f0700000002:vol-610-4d48-306f","result":true,"contextMap":{},"wait":0,"bypassHostMaintenance":false}
- management server
- volume path is set to new, then set to old when fail
- Bug: Actually, on KVM host, the VM is running with new volume, but the new volume is removed
- Input/Output error, read-only system
- New:
- check volume statistics on destination pool via ScaleIO gateway
- if allocation size = provisioned size, volume migration is done. remove volume on source pool.
- if allocation size is not same to provisioned size, volume migration failes, remove volume on destination pool.
| - management
- Resource [StoragePool:12] is unreachable: Migrate volume failed: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 2520889891420635168 to Host 1 timed out after 2400
- Failed to migrate PowerFlex volume
- volume path is set to new, then set to old when fail (looks good)
- revertBlockCopyVolumeOperations
- No issues found, ALL looks good
changes on agent - it uses "dm.blockCopy" (introduced by Hari)
- when agent is disconnected or restarted
- abort the block copy job by a disconnect hook (VolumeMigrationCancelHook)
- similar as MigrationCancelHook
| same as left | Issue: - one volume in Migrating in DB
- path, iscsi_name, pool_id are info on destination pool
- when mgmt server is restarted, it updates all Migrating volume to Ready
- this is disabled if reconcile is enabled
How to reconcile: - only if vm is Running and volume in Migrating state
- The current pool is destination pool of reconcile command
- check vm state and vm disks via agent
- if volume is Ready on source pool,
- revert path, iscsi_name and pool_id to source pool
- create a dummy volume on dest pool
- remove the dummy volume
- if volume is not found on source pool
- update state to Ready
- create a dummy volume on source pool
- remove the dummy volume
- if vm is running on source pool (check by fullpath) = volume is ready on source pool
- if vm is running on dest pool, update volume state to Ready, and remove volume from source pool
|
5. Migrate Volume of Stopped VM (on Powerflex)
| intermittent failure - CopyCommand: wait until connection is recoved
continuous failure - Resource [StoragePool:11] is unreachable: Volume [{"name":"ROOT-533","uuid":"ff8a53a0-d513-4bfc-8920-352948ea066b"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 8886446489732120778 to Host 1 timed out after 2400]
- 2400 = wait (1200) * 2
- new volume is Expunged
- due to expungeVolumeAsync in copyManagedVolumeCallBack
- Not an issue since the VM is Stopped so no data loss
| - Resource [StoragePool:12] is unreachable: Volume [{"name":"ROOT-531","uuid":"4225516d-52fe-465d-a46d-81f97a2e55b2"}] migration failed due to [com.cloud.utils.exception.CloudRuntimeException: Failed to send command, due to Agent:1, com.cloud.exception.OperationTimedoutException: Commands 3050062847636668430 to Host 1 timed out after 2400].
- new volume is Expunged
- due to expungeVolumeAsync in copyManagedVolumeCallBack
- No issues found, ALL looks good
| same as left | Issue: two volumes in DB - source volume: Migrating state
- dest volume: Creating state
- dest.last_id = source_volume_id
How to reconcile - check volume information via host
- determine state of source and dest volumes
- if source volume is Ready
- update source from Migrating to Ready
- update dest volume from Creating to Destroy
- if dest volume is Ready
- update dest from Creating to Ready
- update source volume from Migrating to Destroy
TO be fixed - when migrate volume from powerflex to powerflex with same system id, it will not create CopyCommand
|
NFS to NFS | NFS to Local | Local to Local | Local to NFS | Powerflex to Powerflex | Powerflex to NFS | NFS to Powerflex | | Migrate VM | MigrateCommand (source) | - | - | - | MigrateCommand | - | - |
Migrate VM with volumes | CopyCommand (template to primary) CreateObjectCommand (new volume) ModifyTargetsCommand PrepareForMigrationCommand (dest) MigrateCommand (source) | Migrate Volume (of Running VM) | Migrate Volume (of Stopped VM)