Status

Current state: Under Discussion

Discussion thread: https://lists.apache.org/thread/z4o5093zqxmc18nzb74vyhvb1hb399v4

JIRA:

Please keep the discussion on the mailing list rather than commenting on the wiki (wiki discussions get unwieldy fast).

Motivation

The Kafka consumer allows users to pause()  and resume()  individual partitions. When a partition is paused, subsequent calls to poll()  will not return any records from that partition until it is resumed. This mechanism is used for application-level flow control and backpressure management.

Higher-level frameworks like Kafka Streams make extensive use of pause/resume internally:

Currently, there are no consumer metrics that expose pause/resume state. Operators and developers have no way to answer basic questions through monitoring:

  1. How many partitions are currently paused? A number of paused partitions may indicate backpressure issues, slow state restoration, or application-level processing bottlenecks.
  2. Which specific partitions are paused? Knowing which partitions are paused helps pinpoint which topics or partitions are experiencing issues.
  3. How long has a partition been paused? A partition that has been paused for an extended period may indicate a stuck consumer, an unrecoverable error in processing, or a state store restoration that is taking too long.

This KIP proposes adding new consumer metrics that track paused partitions and their pause duration.

Public Interfaces

New metrics

consumer-coordinator-metrics

The metric will have the following tags:

Metric NameTypeDescription
paused-partitions-countGauge (Integer)The current number of partitions that have been paused by the user via Consumer.pause() .
paused-partitions-rateRateThe per-second rate of Consumer.pause()  calls.
paused-partitions-totalCumulativeCountThe total cumulative number of Consumer.pause()  calls.

consumer-fetch-manager-metrics

Each metric will have the following tags:

Metric NameTypeDescription
paused-partitionsGauge (Integer)Whether this partition is currently paused. Returns 1  if paused, 0  if not paused.
paused-partitions-duration-secondsGauge (Long)The time in seconds since this partition was paused. Returns -1  if the partition is not currently paused.
paused-partitions-rateRateThe per-second rate of Consumer.pause()  calls for this partition.
paused-partitions-totalCumulativeCountThe total cumulative number of Consumer.pause()  calls for this partition.

Note: These are per-partition metrics. Operators should be aware that registering metrics per partition increases cardinality linearly with the number of assigned partitions. In environments with large partition counts, this may increase memory usage in the metrics registry and in downstream monitoring systems. The paused-partitions-count aggregate metric is available at INFO level for general monitoring, while the per-partition metrics are intended for detailed debugging when partition-level visibility is needed.

Proposed Changes

Track Pause Timestamp

Currently, each partition's internal state tracks whether it is paused with a simple boolean flag. To support the paused-partitions-duration-seconds  metric, this KIP adds a timestamp that records when the partition was paused. The timestamp is set when pause()  is called, cleared (reset to -1) when resume() is called, and also reset to -1 on partition reassignment regardless of prior pause status. This ensures that a partition re-assigned to the same consumer after a rebalance starts with a clean pause state.

consumer-coordinator-metrics

A paused-partitions-count  gauge is registered in the consumer's top-level metrics. When queried, this gauge computes the current count of paused partitions by reading from the consumer's subscription state. Since the consumer already maintains the set of paused partitions internally, no additional data structures are needed for this metric.

consumer-fetch-manager-metrics

The paused-partitions  and paused-partitions-duration-seconds  gauges are registered per partition in the fetch manager metrics, alongside the existing per-partition lag and lead metrics. They follow the same lifecycle:

Compatibility, Deprecation, and Migration Plan

This change only adds new metrics. No existing metrics or APIs are deprecated.

Test Plan

Rejected Alternatives

N/A