DUE TO SPAM, SIGN-UP IS DISABLED. Goto Selfserve wiki signup and request an account.
Status
Current state: Discussion
Discussion thread: here
Vote thread: here
JIRA: KAFKA-19893
PR: https://github.com/apache/kafka/pull/20913
Motivation
This KIP proposes an optional, topic-level feature that provides an opportunity for cost savings on remote storage when the scenario is suitable to use.
Currently, Kafka's tiered storage implementation uploads all non-active local log segments to remote storage immediately, even when they are still within the local retention period.
This results in redundant storage of the same data in both local and remote tiers.
We can see the one redundancy example in the following picture. (just highlight one segement.)
When there is no requirement for real-time analytics or immediate consumption based on remote storage. It has the following drawbacks:
1. Wastes storage capacity and costs: The same data is stored twice during the local retention window
2. Provides no immediate benefit: During the local retention period, reads prioritize local data, making the remote copy unnecessary
Example scenario and this KIP's goal:
Consider a topic with remote storage enable:
- Local retention: 1 day (24 hours)
- Remote retention: 3 days (72 hours)
Data is stored in both tiers for the first day, resulting in about 16 hours of redundant storage. This can leads to cost waste.
So we can reduce the tiered storage redundancy for cost saving. In this case, it will save about 25% cost payment for total size of disk.
We can take an AWS S3 billing example (last 3 months) to show the cost saving detail:
AWS S3 has two cost items (normally, Kafka and the bucket are placed in the same region so that we can ignore the network transfer cost).
The cost item TimedStorage-ByteHrs represents the storage part. In above case, each remote segment lives for about 64 hours. thus, with the new optional solution, each will live for about 48 hours, resulting in approximately 25% cost savings.
In this billing example, the cost would decrease from 67K per quarter. BTW: if our topic is local:1 day + remote: 7 days . the saving part will be about 10% (27K)
Note: You can also refer to Amazon S3 Cost to know more information.
However, this optimization is offered as a topic level optional configuration rather than the default behavior based on followed scenarios:
(1) Some users/topics rely on remote storage for real-time analytics and need the latest data to be available as soon as possible (In fact, it only tries to stay as up-to-date as possible,
but it still can’t include the latest data because the active segment hasn’t been uploaded yet.).
(2) Some topics may set a very high ratio for remote-to-local retention time. The cost savings amount will be small, so it is mainly to avoid waste. Users may think it is not worth enabling the feature for the topics
(3) Kafka admin want to reduce the expansion times for local disk when remote storage down for some time. after all. The already uploaded segments are eligible for deletion from broker when not enable the feature for topic.
You can check the follow picture to understand the logic:
if the remote storage outage for short time: no matter if you enable the feature. it don't have difference.
if the remote storage outage for a medium-term time: you will need one extra expansion: the max size is the your saving cost's part.
if the remote storage outage for a long time: no matter if you enable the feature. you should keep do expansion.
Public Interfaces
This KIP introduces one new topic configuration item: remote.log.keep.latest
BTW: topic's remote storage already had some others items such as remote.log.delete.on.disable/remote.log.copy.disable, etc.
public static final String REMOTE_LOG_LATEST_ENABLE_CONFIG = "remote.log.latest.enable";
public static final String REMOTE_LOG_LATEST_ENABLE_DOC = "Determines whether to upload all segments to remote storage including the latest ones within local retention. " +
"When set to true (default), all committed segments will be uploaded without checking local retention constraints. " +
"When set to false, only segments beyond local retention period will be uploaded to remote storage.";
The default value is true so that the whole remote storage module keeps the original behavior when topic don't set it to false.
Proposed Changes
You can refer to https://github.com/apache/kafka/pull/20913 for the detailed changes.
We change the RemoteLogManager.RLMCopyTask#candidateLogSegments's logic for decide one segment if need to upload to remote:
You can see the uploading will be delayed if the configure remoteLogKeepLatest is false. And After the change, the remote tiered storage redundancy will be reduced with delayed upload.
You can refer to the test case and result: https://github.com/apache/kafka/pull/20913#issuecomment-3547156286
BTW: Here are some additional thoughts/considerations.
- Local files won’t be deleted until they’ve been uploaded to the remote storage, so this change is very safe
You don’t need to worry about files being cleaned up before they be upload to the remote. - Considering the latency of remote storage, the local retention period won’t be set too short.
For example, in our production environment, we keep one day of local data alongside 3-7 days in remote storage, so there’s still one day of redundancy.
Compatibility, Deprecation, and Migration Plan
Backward Compatibility
This change is fully compatible:
Backward compatible: Default value (true) maintains current behavior
Forward compatible: Older clients unaware of this config will ignore it- Deprecation
N/A - Migration for Existing Deployments
N/A. The feature is optional with topic level.
Test Plan
We can use follow tests to cover the change:
Unit Tests:
- Test upload eligibility logic with delay enabled/disabled
- Test configuration validation for topic
Integration Tests:
- Verify segments are uploaded before local segment deleted
- Test the remote storage reduced after topic enable the feature.
Rejected Alternatives
Alternative 1: Make this the default behavior
Reason for rejection: Some users require real-time remote analytics and need data uploaded as soon as possible. Breaking their use case would be unacceptable. after all it is the default behavior before this change.
Alternative 2: Global broker-level configuration only
Reason for rejection: Different topics have different requirements. Topic-level granularity is essential.





