DUE TO SPAM, SIGN-UP IS DISABLED. Goto Selfserve wiki signup and request an account.
...
Current state: ["Under Discussion"]
Discussion thread: here
JIRA: KAFKA-15777
Please keep the discussion on the mailing list rather than commenting on the wiki (wiki discussions get unwieldy fast).
...
When remote storage is enabled on a topic, and the consumer is doing backfill, then the data is read from the remote storage. Currently, Kafka remote storage supports reading data only reading from the first remote partition in the given FETCH request (see KAFKA-14915). The default value of max.partition.fetch.bytes is configured to 1 MB. Reading one 1 MB of data per FETCH request from remote, increases the round-trip time for the consumer and the cloud storages provider connectors are often tuned to read data in chunks of 4 MB.
...
- Introduce new
remote.max.partition.fetch.bytesconsumer config, with default value to 1 MB. If not configured, fallback to usemax.partition.fetch.bytesvalue. - Bump FetchRequest to v18 to add the
RemotePartitionMaxBytesfield to propagate the value configured on the consumer to the broker. Binary log format
The network protocol and api behavior
Any class in the public packages under clientsConfiguration, especially client configuration
org/apache/kafka/common/serialization
org/apache/kafka/commonorg/apache/kafka/common/errors
org/apache/kafka/clients/producer
org/apache/kafka/clients/consumer (eventually, once stable)
Monitoring
Command line tools and arguments
- Anything else that will likely break existing users in some way when they upgrade
Proposed Changes
The proposal is to introduce a new Consumer config: remote.max.partition.fetch.bytes to configure the number of bytes returned from remote storage in a FETCH request. The reason for not re-using the existing the max.partition.fetch.bytes is that:
- The default value of
fetch.max.bytesis set to 50 MB andmax.partition.fetch.bytesis 1 MB. - If the user increases the
max.partition.fetch.bytesvalue to 4 MB, then it applies for all the partitions in the FETCH requests. (ie) Reading the data from local storage too. - Assume that the consumer is reading data from a topic with 64 partitions and 16 partition leaders are co-located in a single broker.
- Broker allocates 12 instances of 4 MB buffers (12 x 4 = 50 MB) which might trigger the Young Gen and Old Gen GC when there are many FETCH requests from different consumers.
- Broker does not use BufferPool to allocate the buffers.
- Allocating big byte-buffers have direct impact on GCs.
- Remote read gets triggered for only one partition in a given FETCH request. And, we want to allocate the big byte-buffer (4 MB) only for remote read requests.
Compatibility, Deprecation, and Migration Plan
- What impact (if any) will there be on existing users?
- If we are changing behavior how will we phase out the older behavior?
- If we need special migration tools, describe them here.
- When will we remove the existing behavior?
Test Plan
- The assumption made is that there will be fewer RemoteRead requests compared to the LocalRead requests. Kafka performs faster when reading the data from PageCache instead of disk.
- This will also simplify the RemoteStorageManager plugin implementation.
| Code Block | ||||
|---|---|---|---|---|
| ||||
{ "name": "Topics", "type": "[]FetchTopic", "versions": "0+",
"about": "The topics to fetch.", "fields": [
{ "name": "Topic", "type": "string", "versions": "0-12", "entityType": "topicName", "ignorable": true,
"about": "The name of the topic to fetch." },
{ "name": "TopicId", "type": "uuid", "versions": "13+", "ignorable": true, "about": "The unique topic ID."},
{ "name": "Partitions", "type": "[]FetchPartition", "versions": "0+",
"about": "The partitions to fetch.", "fields": [
{ "name": "Partition", "type": "int32", "versions": "0+",
"about": "The partition index." },
{ "name": "CurrentLeaderEpoch", "type": "int32", "versions": "9+", "default": "-1", "ignorable": true,
"about": "The current leader epoch of the partition." },
{ "name": "FetchOffset", "type": "int64", "versions": "0+",
"about": "The message offset." },
{ "name": "LastFetchedEpoch", "type": "int32", "versions": "12+", "default": "-1", "ignorable": false,
"about": "The epoch of the last fetched record or -1 if there is none."},
{ "name": "LogStartOffset", "type": "int64", "versions": "5+", "default": "-1", "ignorable": true,
"about": "The earliest available offset of the follower replica. The field is only used when the request is sent by the follower."},
{ "name": "PartitionMaxBytes", "type": "int32", "versions": "0+",
"about": "The maximum bytes to fetch from this partition. See KIP-74 for cases where this limit may not be honored." },
{ "name": "ReplicaDirectoryId", "type": "uuid", "versions": "17+", "taggedVersions": "17+", "tag": 0, "ignorable": true,
"about": "The directory id of the follower fetching." },
// New field added to FetchTopic
{ "name": "RemotePartitionMaxBytes", "type": "int32", "versions": "18+", "ignorable": true,
"about": "The maximum bytes to fetch from this partition from remote storage. See KIP-74 for cases where this limit may not be honored." }
]}]}, |
If the remote.max.partition.fetch.bytes is not configured / FETCH request is from the old client, then it fallbacks to the use the max.partition.fetch.bytes config.
Compatibility, Deprecation, and Migration Plan
The change is fully compatible with the existing code. There won't be any impact to the existing users.
Test Plan
The patch will be covered with unit test and integration tests.Describe in few sentences how the KIP will be tested. We are mostly interested in system tests (since unit-tests are specific to implementation details). How will we know that the implementation works as expected? How will we know nothing broke?
Rejected Alternatives
If there are alternative ways of accomplishing the same thing, what were they? The purpose of this section is to motivate why the design is the way it is and not some other way.