DUE TO SPAM, SIGN-UP IS DISABLED. Goto Selfserve wiki signup and request an account.

DUE TO SPAM, SIGN-UP IS DISABLED. Goto Selfserve wiki signup and request an account.
When a long-running command or job is tasked to the KVM agent, and the agent begins processing it, meanwhile if the agent restarts or crashes, management server assumes it as a job failure. Consequently, the management server attempts to revert the resource to its original state, which might not be accurate since the agent could have already completed the operation.
The FR primarily requires that cases of long running jobs and unknown successes are addressed, when the CloudStack management server loses communication with the agent.
Currently KVM agent does not have any state persistence mechanism for the job commands it processes. If the KVM agent crashes or restarts during processing any job commands it has no way to track such commands to take further action on that. If ongoing job commands can be tracked, the agent can recover or restart and handle cases of silent successes. This functionality can also be used by the KVM agent to inform the management server about command process interruptions. As a result, the management server can attempt to reconcile the resources and their states.
In CloudStack, the management server and KVM agent interact using a Command and Answer pattern. When an operation like volume copy between two storages needs to be performed, the management server issues a Command to the KVM agent. The agent receives this command, processes it by executing the necessary actions on the host, and then formulates an Answer containing the result. Management server waits for the Answer which will be identified by a sequence ID which is unique to the each Command. If any interruption occurs on the agent (like agent restart, agent crash), currently it is not able to send the Answer informing the management server about the Command processing state. Management server assumes it as a job failure on timeout without trying to know the actual resource state
Following are few such common long-running operations where an interruption on the ongoing Command will have an effect on the original resource state and these will be addressed as part of this FR
The following major scenarios are considered for the interruption of on-going command(s) by the KVM agent:
We propose to handle the above scenarios, to reconcile the resource state properly in the following sections.
In the KVM agent whenever it looses connectivity to the management server (Scenario 1) or upon agent restart (Scenario 3), a disconnect event can be generated and during that event processing all ongoing Commands can be handled without having to leave them in unknown state. Each Command can be implemented to register a hook to handle any interupptions to that Command and whenever KVM agent gets a disconnect event, all the registered hooks can be executed to handle that interruption.
For example, during volume copy if the KVM agent gets a disconnect event then the corresponding command wrapper’s disconnect hook will be initiated where in it can cancel that copy operation. The disconnect hook feature (PoC PR: https://github.com/apache/cloudstack/pull/5552/) can be explored to handle any ongoing commands in the agent when agent receives any disconnect event. This adds a way for KVM agent CommandWrappers to register callback code that handles in case of an agent disconnect event.
This adds a way for KVM agent CommandWrappers to register code to run in case of an Agent disconnect event. LibvirtComputingResource with a list to store DisconnectHooks, and then it implements the disconnected() method to call these hooks in the event of a loss of connectivity. More details can be found in this PoC PR here: https://github.com/apache/cloudstack/pull/5552
A new ReconcileResourceManager is proposed that runs within the management server to track resources like volumes, VMs, and snapshots; and activates when a command operation times out or fails.
For example, during volume migration in the KVM agent, if the Command operation is interrupted by an agent connection failure to the management server (Scenario 1) or upon agent crash (Scenario 2) then the management server can handle such events upon Command timed out. When the management server considers it as a command failure, it can use the ReconcileResourceManager to reconcile the state of the volume and persist in the database.
When a command is in progress on the agent and an interruption occurs, the management server will receive an interrupt notification via the Answer. As described, the "ReconcileResourceManager" will be implemented with required methods to handle these interruptions and reconcile the resource. If the original agent is unavailable, the manager will attempt to send new commands to other existing agents to locate the resource and determine its status.
If the KVM agent is restarted while a Command is being processed, the actual operation may be successful but since the management server won’t receive any Answer the operation may eventually time out and treats it as a failure.
To handle the KVM agent restarts (Scenario 3) while processing long running commands, state of commands can be logged as checkpoints by the management server to retain known state information after crash or reconnection.
The command state logging saves the check points while processing the command, for example “START”, “PROCESSING”, “COMPLETED” with all the relevant details of the command. After agent comes up, agent can read the state of the command and take further action. For example, if the state of a command is “PROCESSING” then the agent can inform the management server to reconcile and process it.
In the agent during the command processing in the wrappers, the progress of the command execution will be logged and used after the agent restart to take actions accordingly based on the state of the command’s state.
The command state can be logged as a JSON file within the agent. For example, a CopyCommand might look like this:
{
“commandId”: “1-4192288303128511738”,
“state”: “PROCESSING”,
“timestamp”: “2024-07-25T10:20:30Z”,
“starttime”: “2024-07-25T10:10:30Z”,
”timeout”: “60”,
“request”: “Request:Seq 1-4192288303128511738: { Cmd, MgmtId: 2484172157385, via: 1, Ver: v1, Flags: 100011, [{\"org.apache.cloudstack.storage.command.CopyCommand\":{\"srcTO\".…. “
}
After the agent restarts, a new "CommandInvestigator" thread will process these JSON files. If an incomplete command is found, an Answer will be created with a "commandStatus" parameter indicating the interruption, informing the management server. The JSON files will be saved at “/usr/share/cloudstack-agent/tmp” and deleted after the Answer is sent. All logged commands must be processed, so files will only be deleted after processing. Commands exceeding their timeout will be deleted on agent startup.
For example if the copy command is send to the KVM agent and agent got restarted, following is the workflow that will happen with the new changes to inform the management server that the command has been interrupted
Following the sequence diagram for the above scenario:
| Scenarios | Current behavior | How to deal with it |
|---|---|---|
| Agent cannot send Answer to management server, therefore timed out on management server | Agent: (for reconcile command only)
Management server: (for reconcile command only)
|
| Agent interrupt the process or processing in backend) |
|
| Agent interrupt the process or processing in backend) |
|
| timed out on management server | |
| Agent process the command
|
Management server: CREATED
Agent (all good): STARTED → PROCESSING → COMPLETED → Send Answer to management server → Mgmt server removes the reconcile command jobs (new table: reconcile_command_jobs ?) → Agent removes the JSON file if job is not found.
Agent (wrong): STARTED → PROCESSING → FAILED → Send Answer to management server → Mgmt server removes the reconcile command jobs (new table: reconcile_command_jobs ?) → Agent removes the JSON file if job is not found.
Agent is restarted: STARTED → PROCESSING → INTERRUPTED (processed by java) → Send CommandInfo to management server via PingCommand every minute
Agent is restarted: STARTED → PROCESSING_IN_BACKEND → DANGLED_IN_BACKEND (processed in backend)→ Send CommandInfo to management server via PingCommand every minute
Agent timed out: STARTED → PROCESSING → TIMEDOUT (processed by java)→ Send to Answer management server
Agent timed out: STARTED → PROCESSING_IN_BACKEND → DANGLED_IN_BACKEND (processed in backend) → Send Answer to management server and Send CommandInfo via PingCommand every minute
Management server (TIMEDOUT) → Add CommandInfo to reconcile list
Management server (Timeout) → Add CommandInfo to reconcile list
Agent → First Management server → (If it is not the source mgmt server and not found in the memory, send to) Other management server to reconcile.
API, database, service layer and agent changes
No.
Command.java
public enum State {
CREATED, // Command is created and sent to agent
STARTED, // Command is started by agent
PROCESSING, // Processing by agent
PROCESSING_IN_BACKEND, // Processing in backend by agent
COMPLETED, // Operation succeeds
FAILED, // Operation fails
TIMED_OUT, // Timed out on management server or agent
INTERRUPTED,// Interrupted by agent (for example agent is restarted),
DANGLED_IN_BACKEND // Backend process which cannot be processed normally (for example agent is restarted)
}
new method
public boolean isReconcile() {
return false;
}