DUE TO SPAM, SIGN-UP IS DISABLED. Goto Selfserve wiki signup and request an account.
...
Today, all Kafka protocol requests include the client ID. This identifier can be set by the user by setting a the client configuration property client.id. This is a useful capability, but it has limitations. First, the identifier is sent on each and every request. Users sometimes have use quite long client IDs, even encoding metadata such as the availability zone into the string. The client ID is then sent on every request, in spite of the fact that the Kafka protocol is connection-oriented and it is really only necessary to send the string on the first request after connection initiation. It's conceivable that a client could mutate the client ID between requests but it is expected to be static for the duration of a connection, so sending it repeatedly is wasteful. Second, the identifier client ID is not sufficient to identify a particular client because it is unusual for users to assign unique identifiers to their clients.
This KIP proposes introducing adds a UUID called the client instance ID into the request header of all Kafka protocol requests. Correlating Each client has a different client instance ID, so correlating requests from a particular client becomes much easier as a result.
Apart from its use in client telemetry, the addition of the client instance ID has no significance to the broker. It is being added to improve traceability and problem determination.
Proposed Changes
This KIP proposes extending the scope of two changes to the Kafka protocol.
Client Instance ID
This KIP proposes adding a UUID called the client instance ID concept from KIP-714 into into the request header of all Kafka protocol requests. Those familiar with KIP-714 will be aware that it already introduced the client instance ID, so this KIP actually proposes extending its scope to a universal unique identifier for client instances in all RPCs.
The client instance ID is now calculated by the client during the constructor of the client before it makes its initial connection to the cluster. This differs slightly from how KIP-714 initialized the client instance ID, but the change is compatible with the existing behavior of the broker. The client uses the same client instance ID for its connections to every broker throughout its lifetime, even when it rebootstraps and makes new connections. The consistency and uniqueness of the identifiers are both important characteristics.
The client instance ID is added is introduced as a tagged field in the RPC request header, so it is present on all RPCs which that use the v2 request header, which is almost all of the current versions of the RPCs . For those which do not, the client instance ID will be available in the broker's connection context.The client instance ID is calculated by the client during the constructor of the client before it makes its initial connection to the cluster. The client instance ID generated is never the zero UUID(only SaslHandshake and OffsetDelete use the v1 request header for their latest versions). The client is required to be consistent in its use of client instance ID in request headers. If it specifies the value in its request headers, every request on a connection must specify the same value. If it does not specify the value in its request headers, every request on a connection must not specified a value.
If a connection specifies a client instance ID in the request header of its first request which uses the v2 request header, it must specify the same client instance ID in the request header for all subsequent requests. The initial client instance ID for each connection will be cached by the broker for checking (this is an implementation detail, but caching it in the ChannelMetadataRegistry is an option). Once a client has specified a client instance ID in the request header of its first request, any subsequent requests which are missing the client instance ID or which specify a different value for the client instance ID will be rejected with error code INVALID_REQUEST.
If a client does not specify a client instance ID in the request header of its first request which uses the v2 request header, it must not specify a client instance ID in the request header of any subsequent requests. If it does so, the request will be rejected with error code INVALID_REQUEST.
In KIP-714, the client instance ID is was created by the broker which responds responded to a client's first GetTelemetrySubscription GetTelemetrySubscriptions RPC. In KIP-848, initially the member ID was created by the brokergroup coordinator in response to a heartbeat, but subsequently KIP-1082 changed this so that the client creates created its own member ID. As a result, this KIP also changes the definition of the client instance ID so that the client creates its own UUID before it makes its initial connection to the cluster, and then it uses it for all future requests made to all brokers by that client instance. In addition. After this KIP, when the client makes its first GetTelemetrySubscription GetTelemetrySubscriptions request, it will also supply supplies the client instance ID which it created. There is no need to change the behaviour of the cluster in handling client telemetry requests, it will just be the case that the has already created rather than specifying a zero client instance ID; the Apache Kafka Java client after this KIP no longer supplies a zero client instance ID expecting the broker to calculate the ID. However, for migration purposes, the on GetTelemetrySubscriptions requests. The original behavior of KIP-714 is still supported for clients which do not yet send ClientInstanceId in the request header.Apart from its use in client telemetry, the addition of the client instance ID has no significance to the broker. It is being added to improve traceability and problem determination.; they just specify a zero client instance ID in their first GetTelemetrySubscriptions request and then receive a client instance ID to use for client telemetry calculated by the broker in its response.
Client ID
This KIP also proposes sending a null client ID for all requests on each client connection, with the exception of a second change to the Kafka protocol. It proposes sending the client ID only on the initial request on each connection, and then sending a null client ID for all subsequent requests. This eliminates some redundant the unnecessary overhead of repeatedly sending the same client ID string to the broker on every request. The initial client ID for each connection will be cached by the broker (this is an implementation detail, but caching it in the ChannelMetadataRegistry is an option). After this KIP, the broker will assume that the client ID from the initial request applies to all subsequent requests on a connection, and it will ignore the client ID specified on any subsequent requests.
Public Interfaces
Client API Changes
The client instance ID will be is calculated during the constructor of the Producer , Consumer , ShareConsumer and Admin implementations so there is no need to have a timeout parameter on the accessor method. The following method will be is added to these interfaces:
...
and then the following method will be is deprecated for removal in Apache Kafka 5.0:
...
In a similar vein, the following method in KafkaStreams will be is added:
public ClientInstanceIds clientInstanceIds()
and then the following method will be is deprecated for removal in Apache Kafka 5.0:
...
| Code Block |
|---|
{
"type": "header",
"name": "RequestHeader",
// Version 0 was removed in Apache Kafka 4.0, Version 1 is the new baseline.
//
// Version 0 of the RequestHeader is only used by v0 of ControlledShutdownRequest.
//
// Version 1 is the first version with ClientId.
//
// Version 2 is the first flexible version.
// Tagged Tagfield 0 introduces client instance ID. (KIP-1313).
"validVersions": "1-2",
"flexibleVersions": "2+",
"fields": [
{ "name": "RequestApiKey", "type": "int16", "versions": "0+",
"about": "The API key of this request." },
{ "name": "RequestApiVersion", "type": "int16", "versions": "0+",
"about": "The API version of this request." },
{ "name": "CorrelationId", "type": "int32", "versions": "0+",
"about": "The correlation ID of this request." },
// The ClientId string must be serialized with the old-style two-byte length prefix.
// The reason is that older brokers must be able to read the request header for any
// ApiVersionsRequest, even if it is from a newer version.
// Since the client is sending the ApiVersionsRequest in order to discover what
// versions are supported, the client does not know the best version to use.
{ "name": "ClientId", "type": "string", "versions": "1+", "nullableVersions": "1+", "flexibleVersions": "none",
"about": "The client ID string." },
{ "name": "ClientInstanceId", "type": "uuid", "versions": "2+", "taggedVersions": "2+", "tag": 0, "ignorable": "true",
"about": "UniqueThe unique idID for this client instance." }
]
}
|
...
After this KIP, the Apache Kafka Java client will send the ClientInstanceId in the request header and also in the request body. If present in the request header, it must match the value in the request body, and that value will not be zero, or else the request will be rejected with error code INVALID_REQUEST.
It is still permitted to send the value zero for the ClientInstanceId in the request, which will cause the broker to create a ClientInstanceId and send it back in the response. This is the original KIP-714 behavior and it is still supported for clients which do not use the ClientInstanceId in the request header.
...
After this KIP, the Apache Kafka Java client will send the ClientInstanceId in the request header and also in the request body. If present in the request header, it must match the value in the request body, and that value will not be zero, or else the request will be rejected with error code INVALID_REQUEST.
Compatibility, Deprecation, and Migration Plan
The addition of a tagged field in the request headers header should have no impact.
The removal of broker will assume that the client ID is more perhaps a more significant change. However, the client ID can easily be cached as part of the connection context in the broker. If further research reveals that the client ID removal is more problematic, the KIP can be reduced to just the addition of the client instance IDfrom the initial request on a connection is used on all subsequent requests. There is no need to send a client ID on any requests except the initial request on a connection, and the client ID specified on requests after the initial request on a connection will be ignored. The client ID is expected to be static and not change from request to request, so this change should have no effect.
Test Plan
Unit tests will be added to ensure that the new behaviour works behavior works as expected. The existing integration and system tests should be entirely unaffected by the change, which would show that there was no behavioural behavioral impact.
Rejected Alternatives
It would be possible to add an untagged field to the request header and bump the version of the request header but this is expensive. Each version of the Kafka protocol RPCs has an associated request header version, so it would be necessary to bump the versions of all the other RPCs.