You are viewing an old version of this page. View the current version.

Compare with Current View Page History

Version 1 Next »

Background

VM life cycle in CloudStack is current represented through a number of lifecycle VM states, following is a complete list of these states

 Starting, Running, Stopping, Stopped, Destroyed, Migrating, Expunging, Error, Unknown

Compared with VM states defined in underlying hypervisors, CloudStack lifecycle VM states contains more information the reflect VM's cloud environment, when we say a CloudStack VM is running, it usually means that

  • VM's storage volumes are ready
  • VM's network environment in Cloud is ready
  • VM is powered on at hypervisor host

To manage a CloudStack VM properly, current CloudStack is designed to have hypervisor resource-agent to participate VM state management and periodically sync-back with CloudStack management server. Therefore, in addition for hypervisor resource agent to be aware of hypervisor specific VM power state, hypervisor resource needs also to know about VM's state in CloudStack, especially to those transitional CloudStack VM states like Starting, Migrating, etc.

Upon hypervisor host-connect event, hypervisor host resource-agent will first report all VMs on the host to management server, it triggers a "full-sync" process with management server to build an initial sync start point, the host won't be considered as in UP state until this "full-sync" process is completed. After host is connected, hypervisor host resource agent will periodically perform "delta-sync" with CloudStack management server.

 "Full-sync" and "delta-sync" are currently forming the foundation of VMSync process. Although it works nicely most of time, as CloudStack is adding support to external managers like vCenter, the state sync scenarios can become hard to handle when out-of-band changes posted from external managers, following use cases sometimes can cause problematic issues in normal CloudStack operating

1) Takes a long time to bring up all hypervisor hosts in a large setup

During management restart, if things fall out of sync, "full-sync" on host connect-phase can trigger a series of chain actions (actions to bring state in sync) that takes a long time to finish

2) Activities from user, from HA process and VMSync process can collide and the resolution of conflicts is hard to cover all scenarios.

3) Hyprvisor resource-agent to participate into CloudStack VM state management has increased the complexity for people to write a new hypervisor support.

This improvement effort is to address these issues, these will help CloudStack to better interage third-party virtualization managers like VMware vCenter

Design

 High-level principals

At very high-level, we try to attack the problem in following areas

1) Hypervisor resource-agent to report raw VM power state only

This is to de-couple resource agent from CloudStack VM lifecycle state management, letting hypervisor resource-agent only carry on hypervisor-specific actions and report hypervisor raw VM state can greatly simplify the coding of hypervisor resource-agent

2) Serialize VM operations

Currently, state transition handling always happens at in-place context, for example, when management server receives hypervisor VM state report, the handling of the report is processed within the context, even if there may be another thread that is handling user request on the same VM. Although we try to coordinate by checking the state of the VM, by simplify failing it with concurrent-access exception.  

In the new design, we will try to serialize activities  to the same VM through job facility, as there always be one active operation is in executing, the state transition logic can be simplified. Take the VM migrating case, as it involves with two hosts, in previous model, with VM state report from different host, we have to handle it carefully as the host report may come at un-predicated order. 

3) Message bus to coordinate with activities

We will try to use a message-bus to co-ordinate different activities within the management server. This facility is different with the existing feature of "Event Bus", the later one is mainly to integrate external systems through persist-able message-queue servers.

Code changes

1) Message Bus facility

public interface MessageBus {

void setMessageSerializer(MessageSerializer messageSerializer);

MessageSerializer getMessageSerializer();


void subscribe(String topic, MessageSubscriber subscriber);

void unsubscribe(String topic, MessageSubscriber subscriber);

void clearAll();

void prune();


void publish(String senderAddress, String topic, PublishScope scope, Object args);

}

 
MessageBus defines the interface of the message bus facility, it implements a simple publish/subscribe pattern, publishers and subscribers can linked by sharing a common topic,  topic can be in hierarchy mode, a subscriber at higher hierarchy mode can receive messages from all topics that are below. 

MessageBusBase

A simple message bus implementation

MessageHandler

annotation for subscriber to specify a message handler

MessageDispatcher

For message subscriber to use to dispatch received messages to annotated message handlers

MessageDetector

To detect interested messages on message bus

2) Job facility

AsyncJobManagerImpl

 Refactor it to decouple the tight link with API jobs, make it generic not only executing async API request jobs but also executing internal VM operating jobs

ApiAsyncJobDispatcher

Dispatch async API request jobs

VmWorkJobDispatcher

dispatch internal async VM operation jobs

VmWorkJobVO

VmWorkJobDao

VmWorkJobDaoImpl

Persist classes for internal VM operation jobs

AsyncJobJournalVO

AsyncJobJournalDao

AsyncJobJournalDaoImpl 

Implements job journal facility, all jobs can now have a persist job journal facility

3) Other refactored classes

VirtualMachineManagerImpl

HighAvailabilityManagerImpl

VirtualMachineGuru

ReservationContext

Hypervisor resource classes

etc.

Schema changes

ALTER TABLE `cloud`.`async_job` DROP COLUMN `session_key`;

ALTER TABLE `cloud`.`async_job` DROP COLUMN `job_cmd_originator`;

ALTER TABLE `cloud`.`async_job` DROP COLUMN `callback_type`;

ALTER TABLE `cloud`.`async_job` DROP COLUMN `callback_address`;



ALTER TABLE `cloud`.`async_job` ADD COLUMN `parent_id` bigint;

ALTER TABLE `cloud`.`async_job` ADD COLUMN `job_type` VARCHAR(32);

ALTER TABLE `cloud`.`async_job` ADD COLUMN `job_dispatcher` VARCHAR(64);

ALTER TABLE `cloud`.`async_job` ADD COLUMN `job_executing_msid` bigint;



ALTER TABLE `cloud`.`vm_instance` ADD COLUMN `power_state` VARCHAR(74) DEFAULT 'PowerUnknown';

ALTER TABLE `cloud`.`vm_instance` ADD COLUMN `power_state_update_time` DATETIME;

ALTER TABLE `cloud`.`vm_instance` ADD COLUMN `power_state_update_count` INT DEFAULT 0;

ALTER TABLE `cloud`.`vm_instance` ADD COLUMN `power_host` bigint unsigned;

ALTER TABLE `cloud`.`vm_instance` ADD CONSTRAINT `fk_vm_instance__power_host` FOREIGN KEY (`power_host`) REFERENCES `cloud`.`host`(`id`);



CREATE TABLE `cloud`.`vm_work_job` (

  `id` bigint unsigned UNIQUE NOT NULL,

  `step` char(32) NOT NULL COMMENT 'state',

  `vm_type` char(32) NOT NULL COMMENT 'type of vm',

  `vm_instance_id` bigint unsigned NOT NULL COMMENT 'vm instance',

  PRIMARY KEY (`id`),

  CONSTRAINT `fk_vm_work_job__instance_id` FOREIGN KEY (`vm_instance_id`) REFERENCES `vm_instance`(`id`) ON DELETE CASCADE,

  INDEX `i_vm_work_job__vm`(`vm_type`, `vm_instance_id`),

  INDEX `i_vm_work_job__step`(`step`)

) ENGINE=InnoDB DEFAULT CHARSET=utf8;



CREATE TABLE `cloud`.`async_job_journal` (

  `id` bigint unsigned NOT NULL AUTO_INCREMENT COMMENT 'id',

  `job_id` bigint unsigned NOT NULL,

  `journal_type` varchar(32),

  `journal_text` varchar(1024) COMMENT 'journal descriptive informaton',

  `journal_obj` varchar(1024) COMMENT 'journal strutural information, JSON encoded object',

  `created` datetime NOT NULL COMMENT 'date created',

  PRIMARY KEY (`id`),

  CONSTRAINT `fk_async_job_journal__job_id` FOREIGN KEY (`job_id`) REFERENCES `async_job`(`id`) ON DELETE CASCADE

) ENGINE=InnoDB DEFAULT CHARSET=utf8;

 

API/UI

This is low-level change that should keep API compatible, UI change is also not mandatory, we can have UI change to take advantage of better job management in the future(i.e. job journal for more descriptive error messages) 



  • No labels