You are viewing an old version of this page. View the current version.

Compare with Current View Page History

« Previous Version 2 Next »

This page is meant as a template for writing a KIP. To create a KIP choose Tools->Copy on this page and modify with your content and replace the heading with the next KIP number and a description of your issue. Replace anything in italics with your own description.

Status

Current state: Under Discussion

Discussion thread: TBD

JIRA: TBD

Please keep the discussion on the mailing list rather than commenting on the wiki (wiki discussions get unwieldy fast).

Motivation

KIP-853: KRaft Controller Membership Changes added support for bootstrapping the KRaft state while KAFKA-13830 - Getting issue details... STATUS added support for bootstrapping the metadata state. In other words, the zero checkpoint (00000000000000000000-0000000000.checkpoint) contains the starting state for KRaft while bootstrap.checkpoint contains the starting state for the cluster metadata. This KIP unifies these two checkpoints by moving the starting metadata from bootstrap.checkpoint to the zero checkpoint.

The main advantage of using the zero checkpoint is that it integrates with the rest of the checkpoint mechanisms like checkpoint loading (RaftClient.Listener#handleLoadSnapshot) and checkpoint deletion introduced in KIP-630: Kafka Raft Snapshot. For example, not deleting the bootstrap.checkpoint has cause issues with Kafka startup logic as documented in KAFKA-19191 - Getting issue details... STATUS .

Public Interfaces

bootstrap.checkpoint

The bootstrap.checkpoint file is created by the kafka-storage tool in the metadata.log.dir. This file will be deprecated. The kafka-storage tool will not create this file any more and will instead write metadata record to the zero checkpoint in the cluster metadata partition.

The kafka-storage tool will delete the bootstrap.checkpoint file if it already existing. This is done to make sure that the metadata state is not contained in both bootstrap.checkpoint and the zero checkpoint.

Proposed Changes

Controller

When the controller handles RaftClient.Listener#handleLoadSnapshot if the snapshot id has an epoch of 0 and an offset of 0, the controller will consider these records as the bootstrapping records. The controller will rewrite bootstrap records to the log if they have been successfully written in the past. If the controller doesn't load a snapshot at epoch 0 and offset 0, the controller will load the bootstrap.checkpoint and rewrite the bootstrapping record to the cluster metadata partition.

Compatibility, Deprecation, and Migration Plan

To be compatible with previous bootstrapping of Kafka, at controller activation the controller can be in the following states:

  1. bootstrap.checkpoint exist with metadata records and the zero checkpoint doesn't exists - In this case the controller will behave as it does today. The controller will be able to identify this case because RaftClient.Listener#handleLoadSnapshot won't ask to load the zero snapshot.
  2. bootstrap.checkpoint exist with metadata records and the zero checkpoint exists but doesn't contain any metadata records - In this case the controller will behave as it does today. The controller will be able to identify this czse because RaftClient.Listener#handleLoadSnapshot will ask to load the zero checkpoint but it will be empty, no metadata records.
  3. bootstrap.checkpoint doesn't exist and the zero checkpoint exists with metadata records - In this case the controller will use the zero checkpoint a

Test Plan

Describe in few sentences how the KIP will be tested. We are mostly interested in system tests (since unit-tests are specific to implementation details). How will we know that the implementation works as expected? How will we know nothing broke?

Rejected Alternatives

If there are alternative ways of accomplishing the same thing, what were they? The purpose of this section is to motivate why the design is the way it is and not some other way.

  • No labels