Versions Compared

Key

  • This line was added.
  • This line was removed.
  • Formatting was changed.


IDIEP-25
Author
Sponsor
Created

 

Status
Status
colourGrey
titleDRAFT


Table of Contents

Motivation

Partition Map Exchange is an internal process crucial for Ignite cluster maintaining consistent view of partitions distribution across all nodes in the cluster. Any event triggering change of partitions distribution triggers PME as well, e.g. server nodes joining and leaving, dynamic caches starts etc.

...

Proposed solution is to stop nodes automatically in known situations in which PME hangs.

Description

The following scenarios should be covered:

1 Non-coordinator nodes not replying with local partition maps

When coordinator on phase 1 detects that particular nodes don't reply in time with their local partition maps, it may decide to forcibly stop these nodes to unblock exchange.

It should be possible for user to configure a minimal number of copies. If stopping these nodes leads to partition loss coordinator checks user-provided policy to make this decision. If user allowed such actions coordinator proceeds with stopping otherwise warning is printed to logs.

If there is no risk of partition loss then coordinator stops nodes without any additional checks.

would cause some partitions to have less copies, coordinator should print warning to logs and keep waiting.

2 Non-coordinator node not applying full partition map

When non-coordinator node (say nodeA) has sent local partition map successfully but doesn't receive full partition map in time it should check status of its exchange on coordinator.

If coordinator informs that exchange has already been finished nodeA should stop itself as its partition map is out-of-date with the rest of the cluster.

3 Coordinator node not sending full partition map to other nodes

When coordinator node detects that it received all local partition maps but didn't send back full partition map it should stop itselfIf coordinator hasn't initiated exchange for a given topology version, non-coordinator nodes should kick it from cluster.

Risks and Assumptions

Proposal requires defining new policy in public API for scenario#1, new protocol should be developed for scenario#2 (non-coordinator node requests status of specific exchange from coordinator).

Tickets

Jira
serverASF JIRA
columnskey,summary,type,created,updated,due,assignee,reporter,priority,status,resolution
maximumIssues20
jqlQueryproject = Ignite AND labels IN (iep-25) ORDER BY status
serverId5aa69414-a9e9-3523-82ec-879b028fb15b