Versions Compared

Key

  • This line was added.
  • This line was removed.
  • Formatting was changed.

...

Last Mirrored Leader Epoch (LMLE): This is the greatest leader epoch that the source cluster recognizes in this destination from the other cluster. Or we can say, below this leader epoch (inclusive), the source cluster and destination cluster is in sync in the perspective of leader epoch.

...

1. Cluster B creates a mirror and starts mirroring from the source cluster A. At this point, because the source cluster stores nothing about the LMLE for this partition, -1 will be returned. The "-1" means the source cluster (A) doesn't recognize any leader epoch in destination cluster (B), so all the logs should be truncated. Later, the cluster B mirrors batches from cluster A, including the leader epoch in the batches.

2. Revere mirroring. Cluster B becomes the source cluster. And before cluster B stops the mirroring from A, it stores the LMLE in the internal topic, which is 1 in this case. When cluster A mirrors from cluster B, it will first query the LMLE from cluster B, and then truncate every record beyond this leader epoch. That is, for cluster A and B, leader epochs for [0, 1] are in sync.

Also, in the cluster B, when it stops the mirror from A, it also bumps the leader epoch to a leader epoch > 1. In this case, it is 2, to make sure the leader epoch is increasing. 

Later, the cluster A mirrors batches from cluster B, including the leader epoch in the batches.

3. Reverse mirroring again. Cluster A stores LMLE 3 in the internal topic, and cluster B queries this LMLE and truncate records beyond leader epoch 3. Besides, the leader epoch in cluster A bumps to 4. Later, the cluster B mirrors batches from cluster A, including the leader epoch in the batches.


Please note that after LMLE truncation, it doesn't mean the log between 2 clusters are in sync. It only means the leader epoch history is in sync. The leader epoch history (i.e. the leader-epoch-checkpoint file in the partition directory) contains the entry for each [leader epoch, start offset]. However, the log might still not converge, yet. In the example below, after cluster B truncates LMLE 3, the offset 4 (in leader epoch 3) still exists, which is inconsistent with the source cluster. In this case, the second log truncation will be triggered by the replication protocol by the last fetched leader epoch comparison. So the offset 4 will be truncated this time, and the logs in these 2 clusters are completely in sync.

...

So, to resolve these cases, we introduce a new "topic-level" config mirror.support.unclean.leader.election . By default it is false. When it is true, during the log truncation for LMLE, we will wait for "all replicas" becomes ISR and completes the log truncation. Again, this is to fulfill the assumption above. So after the log truncation for LMLE, even if there is unclean leader election triggered any time, these log in 2 clusters can still converge successfully using the existing replication protocol.

Failure handling

If there are some replicas cannot catch up with the leader due to that are slow network, disk issue, ... etc while log truncation for LMLE, it will cause the mirroring pending or in FAILED state. In this case, users should manually resolve the issue, or re-assign the replicas into other healthy brokers, and start the mirroring again.

...

Because the new added mirror.support.unclean.leader.election config is a dynamic configuration, it could be enabled in the middle of the mirroring, not at the beginning of the mirroring. If it is enabled at the beginning of the mirroring, all records will guarantee be guaranteed to be consistent between these 2 clusters even if unclean leader election happened. But if it is enabled in the middle of the mirroring, the guarantee becomes:

records after next log truncation for LMLE will be consistent. 

For example, cluster A enables unclean leader election for this topic. And during the log truncation for LMLE 1 at [(2) LMLE 1], the mirror.support.unclean.leader.election is false, and the it this is enabled after [(2) LMLE 1] truncation completed, thenIn this case, we will make sure guarantee that all records after next LMLE truncation, i.e. [(4) LMLE 5], will be consistent in the 2 clusters. 

...

The reason of the guarantee is because the unclean leader truncation in cluster could happen after [(4) LMLE: 5] truncation completed like below image. The leader epoch 5 for offset 2 ~ 5 in the non-ISR log didn't get truncated in [(2) LMLE: 1] process, so it causes inconsistent data before [(4) LMLE: 5] (i.e. offset 5). But after [(4) LMLE: 5], all records (or more specifically, all leader epochs) will be consistent with the cluster B because of log truncation are done in "all replicas".

...