Current state: Under discussion
Discussion thread: here
JIRA:
Please keep the discussion on the mailing list rather than commenting on the wiki (wiki discussions get unwieldy fast).
KIP-932 introduced queueing semantics in Kafka with new group and consumer types: share groups and share consumers. Multiple share consumers in a share group can subscribe to user topics and consume data cooperatively without partition limits.
A share consumer can operate in implicit or explicit mode. In implicit mode, the acknowledgement type is fixed to ACCEPT. In explicit mode, each record must be acknowledged explicitly. The acknowledgement types are:
ACCEPT - indicates the record was consumed successfully, causing it to be acknowledged on the server and never re-delivered to other consumers.
REJECT - indicates the record was not consumed successfully, causing it to be archived on the server and never re-delivered to other consumers.
RELEASE - indicates the record was not consumed successfully, causing it to be moved to available state, making it eligible for re-delivery.
When a share consumer polls a batch of records, the server attaches an acquisition lock timeout task to the batch. If the batch is not acknowledged before the acquisition lock expires, it transitions to available state and becomes eligible for re-delivery (unless delivery count limit is exceeded). The broker cancels and clears the acquisition lock timeout task on ACCEPT or REJECT. On RELEASE state is moved to available and acquisition task is restarted on next delivery attempt.
If a user application using a share consumer in explicit mode processes a record for a long time, it cannot use any acknowledgement type to affect the server acquisition lock timeout without changing the record state. If processing exceeds the timeout duration, the acquisition lock expires and the record is re-delivered, which may be undesirable.
This KIP proposes a new acknowledgement type, RENEW, for explicit mode. It lets a share consumer request the broker to renew the acquisition lock timeout, thereby extending the current delivery attempt without changing the server's record state.
There will be new addition in the org/apache/kafka/clients/consumer/AcknowledgeType.java enum where we will add a new entry RENEW with id 4.
public enum AcknowledgeType {
/** The record was consumed successfully. */
ACCEPT((byte) 1),
/** The record was not consumed successfully. Release it for another delivery attempt. */
RELEASE((byte) 2),
/** The record was not consumed successfully. Reject it and do not release it for another delivery attempt. */
REJECT((byte) 3),
/** Consumer needs more time to process the record. Renew the lease. */
RENEW((byte) 4); // New entry per KIP-1222
...
} |
clients/src/main/resources/common/message/ShareFetchRequest.json
// ShareFetch
{
"apiKey": 78,
"type": "request",
"listeners": ["broker"],
"name": "ShareFetchRequest",
// Version 0 was used for early access of KIP-932 in Apache Kafka 4.0 but removed in Apacke Kafka 4.1.
//
// Version 1 is the initial stable version (KIP-932).
//
// Version 2 supports RENEW ack type
"validVersions": "1-2",
"flexibleVersions": "0+",
"fields": [
{ "name": "GroupId", "type": "string", "versions": "0+", "nullableVersions": "0+", "default": "null", "entityType": "groupId",
"about": "The group identifier." },
{ "name": "MemberId", "type": "string", "versions": "0+", "nullableVersions": "0+",
"about": "The member ID." },
{ "name": "ShareSessionEpoch", "type": "int32", "versions": "0+",
"about": "The current share session epoch: 0 to open a share session; -1 to close it; otherwise increments for consecutive requests." },
{ "name": "MaxWaitMs", "type": "int32", "versions": "0+",
"about": "The maximum time in milliseconds to wait for the response." },
{ "name": "MinBytes", "type": "int32", "versions": "0+",
"about": "The minimum bytes to accumulate in the response." },
{ "name": "MaxBytes", "type": "int32", "versions": "0+", "default": "0x7fffffff",
"about": "The maximum bytes to fetch. See KIP-74 for cases where this limit may not be honored." },
{ "name": "MaxRecords", "type": "int32", "versions": "1+",
"about": "The maximum number of records to fetch. This limit can be exceeded for alignment of batch boundaries." },
{ "name": "BatchSize", "type": "int32", "versions": "1+",
"about": "The optimal number of records for batches of acquired records and acknowledgements." },
{ "name": "Topics", "type": "[]FetchTopic", "versions": "0+",
"about": "The topics to fetch.", "fields": [
{ "name": "TopicId", "type": "uuid", "versions": "0+", "about": "The unique topic ID.", "mapKey": true },
{ "name": "Partitions", "type": "[]FetchPartition", "versions": "0+",
"about": "The partitions to fetch.", "fields": [
{ "name": "PartitionIndex", "type": "int32", "versions": "0+", "mapKey": true,
"about": "The partition index." },
{ "name": "PartitionMaxBytes", "type": "int32", "versions": "0",
"about": "The maximum bytes to fetch from this partition. 0 when only acknowledgement with no fetching is required. See KIP-74 for cases where this limit may not be honored." },
{ "name": "AcknowledgementBatches", "type": "[]AcknowledgementBatch", "versions": "0+",
"about": "Record batches to acknowledge.", "fields": [
{ "name": "FirstOffset", "type": "int64", "versions": "0+",
"about": "First offset of batch of records to acknowledge."},
{ "name": "LastOffset", "type": "int64", "versions": "0+",
"about": "Last offset (inclusive) of batch of records to acknowledge."},
{ "name": "AcknowledgeTypes", "type": "[]int8", "versions": "0+",
"about": "Array of acknowledge types - 0:Gap,1:Accept,2:Release,3:Reject,4:Renew."} // Version 2 supports RENEW ack type (KIP-1222)
]}
]}
]},
{ "name": "ForgottenTopicsData", "type": "[]ForgottenTopic", "versions": "0+",
"about": "The partitions to remove from this share session.", "fields": [
{ "name": "TopicId", "type": "uuid", "versions": "0+", "about": "The unique topic ID."},
{ "name": "Partitions", "type": "[]int32", "versions": "0+",
"about": "The partitions indexes to forget." }
]}
]
} |
clients/src/main/resources/common/message/ShareAcknowledgeRequest.json
// ShareAcknowledge
{
"apiKey": 79,
"type": "request",
"listeners": ["broker"],
"name": "ShareAcknowledgeRequest",
// Version 0 was used for early access of KIP-932 in Apache Kafka 4.0 but removed in Apacke Kafka 4.1.
//
// Version 1 is the initial stable version (KIP-932).
//
// Version 2 will have RENEW ack type
"validVersions": "1-2",
"flexibleVersions": "0+",
"fields": [
{ "name": "GroupId", "type": "string", "versions": "0+", "nullableVersions": "0+", "default": "null", "entityType": "groupId",
"about": "The group identifier." },
{ "name": "MemberId", "type": "string", "versions": "0+", "nullableVersions": "0+",
"about": "The member ID." },
{ "name": "ShareSessionEpoch", "type": "int32", "versions": "0+",
"about": "The current share session epoch: 0 to open a share session; -1 to close it; otherwise increments for consecutive requests." },
{ "name": "Topics", "type": "[]AcknowledgeTopic", "versions": "0+",
"about": "The topics containing records to acknowledge.", "fields": [
{ "name": "TopicId", "type": "uuid", "versions": "0+", "about": "The unique topic ID.", "mapKey": true },
{ "name": "Partitions", "type": "[]AcknowledgePartition", "versions": "0+",
"about": "The partitions containing records to acknowledge.", "fields": [
{ "name": "PartitionIndex", "type": "int32", "versions": "0+", "mapKey": true,
"about": "The partition index." },
{ "name": "AcknowledgementBatches", "type": "[]AcknowledgementBatch", "versions": "0+",
"about": "Record batches to acknowledge.", "fields": [
{ "name": "FirstOffset", "type": "int64", "versions": "0+",
"about": "First offset of batch of records to acknowledge." },
{ "name": "LastOffset", "type": "int64", "versions": "0+",
"about": "Last offset (inclusive) of batch of records to acknowledge." },
{ "name": "AcknowledgeTypes", "type": "[]int8", "versions": "0+",
"about": "Array of acknowledge types - 0:Gap,1:Accept,2:Release,3:Reject,4:Renew" } // Version 2 supports RENEW ack type (KIP-1222)
]}
]}
]}
]
} |
clients/src/main/java/org/apache/kafka/clients/consumer/ShareConsumer.java interface to help the application determine RENEW interval. /** Returns the acquisition lock timeout value in milliseconds for the last set of records fetched from the brokers. */ public Optional<Integer> acquisitionLockTimeoutMs(); |
clients/src/main/java/org/apache/kafka/common/errors/UnsupportedAcknowledgeTypeException.java public class UnsupportedAcknowledgeTypeException extends InvalidConfigurationException {
private static final long serialVersionUID = 1L;
public UnsupportedAcknowledgeTypeException(String message, Throwable cause) {
super(message, cause);
}
public UnsupportedAcknowledgeTypeException(String message) {
super(message);
}
} |
KIP-932 added two new RPCs for record fetching and acknowledging: ShareFetch and ShareAcknowledge. From the share consumer, these RPCs are called when an application invokes poll(Duration)and commitSync(Duration)/commitAsync().
In implicit mode, an application calls poll() on the share consumer and receives a batch of records. It can acknowledge these records on a subsequent call to poll() (piggybacking on ShareFetch). Calling commitSync()/commitAsync() sends a ShareAcknowledge but only the ACCEPT acknowledgement type is used. Calling acknowledge(ConsumerRecord, AcknowledgeType) is illegal in this mode. This KIP does not target implicit mode.
In explicit mode, polling works the same, but the application must acknowledge each record explicitly using one of the AcknowledgeType values. It is here that RENEW comes into play, should the application need it.
If record processing might take time:
AcknowledgeType.RENEW as an argument to acknowledge(ConsumerRecord, AcknowledgeType) for that record.commitSync()/commitAsync() or poll() to send the acknowledgement to the server. Furthermore, the application could also leverage the new ShareConsumer.acquisitionLockTimeoutMs() method to decide on the timeframe in which to make the RENEW call.There is a case where an application tries to send a RENEW acknowledgement to an old broker which does not support the new type. To remedy this, we will also update the ShareConsumer.acknowledge(ConsumerRecord, AcknowledgeType) to throw a new exception UnsupportedAcknowledgeTypeException. This will be handled completely on the client side where the share consumer will use the RPC API version to determine acknowledgement type validity.
When we RENEW a record, the record state for the same maintained by the SharePartition does not change. This means that on subsequent polls by the same member, the record will not be returned by the broker until the lock timeout expires. Due to this, the application might not get any update about the aforementioned record. To remedy this, we propose buffering any RENEW acknowledged records on the share consumer and returning them on subsequent polls until they are re-delivered by the broker at which point the buffer entry can be cleared.
We assume the application uses appropriate concurrency constructs to process records in separate threads.
On the broker side, receiving a RENEW acknowledgement for a specific batch or offset will cancel the existing acquisition lock timeout task and start a new one with the same timeout value as group.share.record.lock.duration.ms. There is one caveat here. If the application causes a ShareFetch RPC to be sent (poll() call) on which RENEW acknowledgements are piggybacked, it could happen that the renewed acquisition lock again times out before the poll() completes. To get around this, we will not return any data from the broker side on ShareFetch requests containing RENEW acknowledgements. That way we are guaranteed that the poll() completes timely.
If the RENEW request received by a broker is invalid and an error is returned, the application should use existing acknowledgement commit callback to listen for the same. Specifically, the application should set a commit response handler callback in ShareConsumer.setAcknowledgementCommitCallback(AcknowledgementCommitCallback) and handle errors, if any. This is the usual mechanism to check commit status in explicit mode and this KIP does not make any change in this code path.
Errors.INVALID_REQUEST to the client. This will also be the case when an old broker receives RENEW request. There is no change required in this path.core/src/test/java/kafka/server/share/SharePartitionTest.java to verify batch and offset level renewal as well as mix acknowledgement type handling.clients/clients-integration-tests/src/test/java/org/apache/kafka/clients/consumer/ShareConsumerTest.java to verify mix and renew acknowledgement type handling.