Versions Compared

Key

  • This line was added.
  • This line was removed.
  • Formatting was changed.

...

However proposed improvements may enable more efficient algorithms for detecting and handling split-brain, e.g. fast detecting of a split-brain on the edges of DCs or simpler and more efficient TopologyValidator implementations based on the notion of "Leader/Follower DC". However all these ideas are a matter of discussion and clarification.

Thin client improvements

Thin client has partition-awareness feature. With this feature client connects to all cluster server nodes and tries to find correct node for key and send request to this node (see IEP-23: Partition Awareness for Thin Clients). To achive this client requests primary node map for cache from server. Non partition-awareness requests (not related to cache requests, or without specified key, which can be mapped to node) distributed randomly to all nodes.

In case of multi DC we can make following improvements to PA mechanism:

  • PA "write" requests still should be sent to primary nodes (nothing changed), but PA "read" requests (if readFromBackup flag is set and write synchrinization mode is not PRIMARY_SYNC) can be sent to backup node, if node is located in the same DC as client (if there is no partitions in the same DC, request should be sent to primary node).
  • Non-PA requests should sent to random node in the same DC as client. 

It's proposed to change "cache partitions" messages (add information about partition in current DC) and introduce new "data center nodes" messages to achive these improvements. Alternatevely, instead of introducing new message we can reuse CLUSTER_GROUP_GET_NODE_IDS message, but in this case client should know about server attribute name to store DC ID (which is server internal information), and everything can be broken if server will change DC ID storage place. 

Protocol changes

Operation codes

The new operation for "data center nodes" request is required:

NameCode
OP_CLUSTER_GET_DC_NODES


5103

OP_CLUSTER_GET_DC_NODES message format

Request
StringData center ID


Response
intNodes count
(UUID) * countNodes IDs

OP_CACHE_PARTITIONS message format changes

Request
boolCustom mapping
StringData center ID  - new field
intCaches count
(int) * countCache IDs


Response
longMajor topology version
int Minor topology version
intPartition mappings count
(Caches configuration and partition mapping) * countCaches configuration and partition mappings


Caches configuration and partition mapping
boolApplicable. Flag that shows, whether standard affinity is used for caches.
intCount of caches
(Cache ID + key configuration) * countCache IDS and cache key configurations
Partition mappingPrimary partition mapping (if Applicable is true)
Partition mappingCurrent DC partition mapping (if Applicable is true) - new field


Partition mapping
intNodes count
(Node ipartitions) * countNode partitions information


Node partitions
UUIDNode ID
intPartitions count
(int) * countPartitions


Risks and Assumptions

Current approach to this IEP introduces components' mostly internal logic modifications, no public API changes or breaking binary compatibility are needed. Protocols stay the same as well with some internal tweaks and possibly some refactoring.

...