Current state: "Under Discussion"
Discussion thread: here
Vote thread: here
JIRA: KAFKA-19784
PR: https://github.com/apache/kafka/pull/20691
Apache Kafka supports rack-aware partition assignment to improve fault tolerance and reduce cross-rack network traffic/cost. However, the current Admin API does not expose rack information for all type of consumer group members, despite this information being available at the protocol level.
The ConsumerGroupDescribeResponse (API Key 69) protocol includes a rackId field for each group member, as defined in the protocol specification:
{ "name": "RackId", "type": "string", "versions": "0+",
"nullableVersions": "0+", "default": "null",
"about": "The member rack ID." }
However, when users call AdminClient.describeConsumerGroups(), the returned MemberDescription objects do not include this rack information. The rack ID is available in the wire protocol but is discarded during response processing in DescribeConsumerGroupsHandler.
The Similar case is for AdminClient.describeShareGroups():
Class | Type of Group | contain RackId? | protocoal response contain rackId? |
MemberDescription | Consumer Group | ❌ | ✅ 是 |
ShareMemberDescription | Share Group | ❌ | ✅ 是 |
StreamsGroupMemberDescription | Streams Group | ✅ | ✅ 是 |
This limitation creates several issues:
Inconsistency: Other group types expose rack information:
StreamsGroupMemberDescription includes rackId() methodTake one example: currently we have to implemen our AZ/Rack analysis using a workaround — passing the rack information into the clientId field and parsing it afterward.
kafkaConsumerConfig.customConfig(ConsumerConfig.CLIENT_ID_CONFIG, generateClientIdWithRack(ip, rack));

Add rack ID support to org.apache.kafka.clients.admin.MemberDescription:
public class MemberDescription {
private final String memberId;
private final Optional<String> groupInstanceId;
private final Optional<String> rackId; // NEW FIELD
private final String clientId;
private final String host;
//omit other codes
}
|
Add rack ID support to org.apache.kafka.clients.admin.MemberDescription:
public class MemberDescription {
private final String memberId;
private final Optional<String> rackId; // NEW FIELD
private final String clientId;
private final String host;
//omit other codes
}
|
public class ShareMemberDescription {
public interface RemoteLogMetadataManager extends BrokerReadyCallback, Configurable, Closeable |
You can refer to https://github.com/apache/kafka/pull/20203/files
We postpone the TopicBasedRemoteLogMetadataManager's initialization part (quering metedata from remote topic) after the server is ready for the request.
Then the retry time is reasonable without worried about different kafka clusters. and the connection not available log won't be seen.
Describe in few sentences how the KIP will be tested. We are mostly interested in system tests (since unit-tests are specific to implementation details). How will we know that the implementation works as expected? How will we know nothing broke?
If there are alternative ways of accomplishing the same thing, what were they? The purpose of this section is to motivate why the design is the way it is and not some other way.