You are viewing an old version of this page. View the current version.

Compare with Current View Page History

« Previous Version 6 Next »

Status

Current state:  Draft

Discussion thread:  here

Vote thread:  here

JIRA

PR:  

Motivation


Consumer group rebalance is one of the most critical lifecycle events for Kafka consumers, because it directly affects users' message consume behavior. At the same time, Kafka is recommending and adopting server-side rebalance.

Kafka clients already provide rebalance-related callbacks, but relying on client-side behavior is operationally fragile. In practice, callback availability and behavior depend on SDK adoption and upgrade cycles, which are difficult to control sometimes. Even when teams implemented the callback with log only. users may reduce or disable the logs by mistake or some other reasons. This creates a common failure mode: when incidents occur, critical rebalance evidence is missing at the client side, making diagnosis slow and uncertain so that we have to query the log in kafka broker side.

However, today server-side observability is limited: coordinator logs can indicate rebalance activity, but logs alone are not a reliable integration surface for automated handling, and existing metrics do not carry the most actionable identifier (consume group id). As a result, operators can observe aggregate rebalance trends, but cannot lightweightly attach custom processing to specific groups without log parsing or invasive client changes.

This KIP proposes a lightweight broker-side rebalance callback capability to expose key rebalance context (including consume group id) on Kafka broker. 

In short, the goal is to provide a stable, centrally controlled extension point for observability and operational automation, without forcing client SDK upgrades and without introducing high-cardinality metric tags. 

Public Interfaces

This KIP introduces one new broker-side extension interface and one new broker configuration.

New Interface

org.apache.kafka.coordinator.group.api.ConsumerGroupRebalanceListener:

 // callback when a consumer group rebalance is detected by the Group Coordinator.

void onConsumerGroupRebalance(String groupId, int groupEpoch, long metadataHash); 

New Broker Configuration

group.consumer.rebalance.listener.classes (LIST, default: empty)

  • A list of fully qualified class names implementing ConsumerGroupRebalanceListener.
  • Classes are instantiated and configured by the broker at startup.
  • Empty value means callback feature is disabled 

Proposed Changes

Add a lightweight broker-side callback path for consumer-group rebalance events, without changing client APIs and without adding high-cardinality metrics.

Server load the group.consumer.rebalance.listener.classes when startuping and trigger the callback interface when rebalance happened.

Refer to the PR example:

Compatibility, Deprecation, and Migration Plan 

  • Backward Compatibility


  • Deprecation
    N/A
  • Migration for Existing Deployments
    N/A. The feature is optional

Test Plan

We can use follow tests to cover the change:

Integration Tests:  

Rejected Alternatives

Alternative 1:  Use the client consume rebalance listener
Relying on client-side behavior is operationally fragile. In practice, callback availability and behavior depend on SDK adoption and upgrade cycles, which are difficult to control sometimes. Even when teams implemented the callback with log only. users may reduce or disable the logs by mistake or some other reasons. This creates a common failure mode: when incidents occur, critical rebalance evidence is missing at the client side, making diagnosis slow and uncertain

Alternative 2:  Enhance the existed server side metric

Enhancing existing coordinator metrics is not sufficient for this use case because the key troubleshooting dimension is group.id, while server metrics are intentionally aggregate. If we add group.id as a metric label/tag, it introduces high-cardinality time series (potentially thousands of groups), which increases memory and storage cost on both broker and monitoring systems and can degrade observability performance

 



  • No labels