Versions Compared

Key

  • This line was added.
  • This line was removed.
  • Formatting was changed.

...

--cluster-id is now optional for brokers + observer controllers. This flag is still required for "bootstrapping" controllers (i.e. controllers who are part of an initial dynamic voter set (determined by the --standalone or --initial-controllers) flags, or who are part of a static voter set).

Proposed Changes

Option 1: Continue to persist cluster id in meta.properties but have KRaft discover it + persist it

  • Rough design:
    • Node can complete a future to allow this value to be discovered by readers outside of kraft layer who need it during startup
    • Raft layer is brought up early during startup, so it is fine to wait until this future completes to proceed with initializing the server
    • Brokers/observers can start Kafka with no cluster id, and rely on the fetch response to discover it in-memory
      • If discovered, node persists cluster id to meta.properties during the startup process before in-memory readers of cluster id
  • Pros:
    • Backwards compatibility is straightforward, since new nodes on old clusters keep using meta.properties for persisting cluster id
    • This functionality is not tied to a MetadataVersion, meaning that any Kafka broker/observer with a software version that supports this KIP can use it, rather than the whole cluster needing to be on some MV >= X.
    • Kraft can easily do its own cluster ID validation for its RPCs, since nodes receive cluster ID via the fetch response if they do not know it and can update that state in-memory + persist it
  • Cons:
    • Since each local node's meta.properties  is its source of truth, if this file is deleted or is changed, it means the broker/observer controller can join another cluster.
      • This is no worse than what exists currently.


Option 2: Introduce a metadata record for cluster id

  • Rough design:
    • Introduce a new Metadata Version that supports a ClusterIDRecord.
    • Brokers/observers can start Kraft KRaft with no cluster id, and rely on metadata publishing pipeline to discover it in-memory
      • Upon discovering the cluster ID for the first time, these nodes need to persist this to meta.properties, and update the raft client in-memory.
      • Ideally, starting KRaft with no cluster id should only be allowed the "first time" (i.e. if the cluster metadata partition doesn't exist on the node yet).
    • Upon restart, nodes need to compare the cluster id from the metadata partition with their local meta.properties . The value in the metadata partition takes precedence.
      • If these values are different, log an error or crash.
      • If cluster ID exists in the metadata partition but not in meta.properties , write it They will persist this cluster ID to meta.properties  too .
    • Bootstrap controllers can add a mandatory “cluster id” record during formatting, or the initial leader can randomly generate a UUID as part of bootstrap metadata records write
  • Pros:
    • Fetch replication automatically handles persistence of the cluster id for each local node
    • Raft module remains independent from metadata module in that KRaft is only responsible for consensus. ClusterID is simply another piece of metadata on which Kraft achieves consensus
  • Cons:
    • Currently, KRaft client also needs to be aware of the cluster ID for its own RPC handling, but the raft module does not decode metadata records
      • Can duplicate the cluster ID as a control record
      • Having a mechanism for “pushing-down” cluster ID from metadata to raft may be complicated.
      • We can duplicate data and have a raft level control record for cluster ID.
      • The fact that the raft client does currently do validation on cluster id does make it unique (i.e. it is used by both metadata and raft layers to prevent nodes from talking to different “clusters”, which is an argument for option 1).
        • For example, if Kraft was used to replicate other data besides the metadata partition, there would still be a concept of cluster id, which needs to be the same across all partitions on the node being managed by Kraft.
    • Users of cluster id must wait until after the node catches up to the metadata LEO and fetches the cluster ID
      • This might be okay, since we block controller server startup on things like authorizer futures completing, which also rely on fetching the metadata log.
      • Controller objects that are initialized with cluster id currently:
        • Authorizer futures: we block on these anyways, can move the construction of `endpointReadyFutures` to after local node fetches cluster ID
        • QuorumController/ClusterControlManager: Currently set to random UUID if DNE. Used to reject broker registration if IDs do not match.
        • ControllerApis: Reported in describe cluster response
        • ControllerRegistrationManager: reads but doesn’t use cluster ID
        • DynamicTopicCLusterQuotaPublisher: this is a publisher, so the handling is pretty trivial, just set cluster ID via onMetadataUpdate
    • Upon restart, in terms of ensuring a node does not talk to another cluster during its lifetime in the presence of deletions/changes to meta.properties , this approach is no better than Option 1 unless the local node can use the cluster id value in its local metadata partition BEFORE contacting the leader and learning of the HWM.
      • The local node does not know if the cluster id record in its metadata log is committed or not until it contacts the leader. However, in order to ensure the local node does not talk to another cluster, it needs to provide a cluster ID in its fetch request to discover the HWM. This is a circularity.
      • In practice, this might be okay, because if ClusterIdRecord with value X was written as part of the bootstrap metadata records write (this is the only way it is possible for a node to have this record in its metadata partition), this means that a KRaft quorum with cluster id X was formed
      Backwards compatibility/migration is non-trivial unless existing clusters are allowed to default to reading `meta.properties` if the cluster ID metadata record does not exist yet.
      • ClusterID record can only be written to the log after a leader with a new version containing this KIP is elected.

Compatibility, Deprecation, and Migration Plan

...

  • Unit tests
  • Integration tests
  • System tests to verify cross-software-version compatibility

Rejected Alternatives

Continue to persist cluster id in meta.properties but have KRaft discover it + persist it

...

  • Node can complete a future to allow this value to be discovered by readers outside of kraft layer who need it during startup
  • Raft layer is brought up early during startup, so it is fine to wait until this future completes to proceed with initializing the server
  • Brokers/observers can start Kafka with no cluster id, and rely on the fetch response to discover it in-memory
    • If discovered, node persists cluster id to meta.properties during the startup process before in-memory readers of cluster id

...

  • Backwards compatibility is straightforward, since new nodes on old clusters keep using meta.properties for persisting cluster id
  • Kraft can easily do its own cluster ID validation for its RPCs, since nodes receive cluster ID via the fetch response if they do not know it and can update that state in-memory + persist it

...

  • Main question to answer is: why not use KRaft itself to bootstrap the cluster id like with other “cluster metadata” (e.g. metadata version)?

...

WIP