DUE TO SPAM, SIGN-UP IS DISABLED. Goto Selfserve wiki signup and request an account.
...
Current state: Under Discussion
Discussion thread: here [Change the link from the KIP proposal email archive to your own email thread]
JIRA: here
Please keep the discussion on the mailing list rather than commenting on the wiki (wiki discussions get unwieldy fast).
Motivation
The log.segment.bytes broker config (and its topic-level synonym segment.bytes) is currently defined as ConfigDef.Type.INT, capping the maximum segment size at Integer.MAX_VALUE (2,147,483,647 bytes, approximately 2 GB). Additionally, the .index file format stores physical file positions as 4-byte signed integers, which also cannot address beyond approximately 2 GB.
With modern storage hardware (multi-TB NVMe drives) and high-throughput workloads, the 2 GB cap is increasingly a problem:
- Excessive file handle usage: Each segment needs 4 files (
.log,.index,.timeindex,.txnindex). A 10 TB partition with 2 GB segments means approximately 20,000 open files. - Frequent segment rolls: A topic ingesting 500 MB/s rolls a new segment every approximately 4 seconds, amplifying index build, flush, and cleaner overhead.
- More log cleaning and compaction work: More segments means more compaction cycles with more small groups.
- Remote storage overhead: Each segment is an individual unit for tiered storage copy/delete operations.
Allowing segments of 4 GB, 8 GB, or larger would significantly reduce these overheads for high-throughput, large-retention workloads.
Public Interfaces
Configuration changes
| Config | Current type | New type | Current range | New range |
|---|---|---|---|---|
log.segment.bytes (broker) | INT | LONG | [1 MB, 2,147,483,647] | [1 MB, Long.MAX_VALUE] (after MetadataVersion finalization) |
segment.bytes (topic) | INT | LONG | [1 MB, 2,147,483,647] | [1 MB, Long.MAX_VALUE] (after MetadataVersion finalization) |
The expanded range (values greater than Integer.MAX_VALUE) is gated by MetadataVersion. Before finalization, the effective range remains [1 MB, Integer.MAX_VALUE]. After IBP_4_4_IV1 is finalized, the range becomes [1 MB, Long.MAX_VALUE].
On-disk format changes
Offset index (.index) file format -- new 12-byte entry format (gated by MetadataVersion):
| Field | Legacy format (8 bytes per entry) | Large format (12 bytes per entry) |
|---|---|---|
| Relative offset | 4-byte signed int | 4-byte signed int |
| Physical position | 4-byte signed int (max approximately 2 GB) | 8-byte signed long (effectively unlimited) |
The large format is only written after MetadataVersion IBP_4_4_IV1 is finalized. Before finalization, all index files use the legacy 8-byte format.
Time index (.timeindex) -- no format change. Entry size is already 12 bytes (8-byte timestamp + 4-byte relative offset). No physical positions are stored.
Transaction index (.txnindex) -- no format change. Uses FileChannel directly with long positions.
Java API changes
| Class | Member | Before | After |
|---|---|---|---|
LogConfig | DEFAULT_SEGMENT_BYTES | int | long |
LogConfig | segmentSize() | returns int | returns long |
LogConfig | initFileSize() | returns int | returns long |
AbstractKafkaConfig | logSegmentBytes() | returns Integer via getInt() | returns Long via getLong() |
RollParams | maxSegmentBytes | int | long |
OffsetIndex | append(long offset, ...) | int position | long position |
OffsetPosition | position field | int | long |
FileRecords | internal size field | AtomicInteger | AtomicLong |
LogSegment | new sizeInBytesLong() | N/A | returns long |
LogSegment | recover() | returns int | returns long |
LogSegment | truncateTo() | returns int | returns long |
LogOffsetMetadata | relativePositionInSegment | int | long |
SegmentPosition (raft) | relativePosition | int | long |
RemoteStorageManager | fetchLogSegment(metadata, int) | only overload | @Deprecated; new default method with long added |
RemoteStorageManager | fetchLogSegment(metadata, int, int) | only overload | @Deprecated; new default method with long, long added |
Not changed
BaseRecords.sizeInBytes()remainsint(449 callers across 89 files -- cascading this change is too large for this KIP).RecordBatch.sizeInBytes()remainsint(bounded bymax.message.bytes).MemoryRecords.sizeInBytes()remainsint(bounded byByteBuffercapacity).- Other coordinator segment configs (
transaction.state.log.segment.bytes,offsets.topic.segment.bytes,share.coordinator.state.topic.segment.bytes) remainINT.
Monitoring
No new metrics are added. Existing segment size metrics will report accurate values for segments larger than 2 GB because LogSegments.sizeInBytes() uses long arithmetic internally.
Command line tools
kafka-log-dirs.sh and DumpLogSegments correctly handle segments larger than 2 GB. DumpLogSegments uses the sliceLong() method for position-based slicing.
Proposed Changes
Phase 1: Config type change
Change log.segment.bytes and segment.bytes from ConfigDef.Type.INT to ConfigDef.Type.LONG. Apply atLeast(1024 * 1024) as the validator. The expanded range (values greater than Integer.MAX_VALUE) is gated by MetadataVersion IBP_4_4_IV1.
Phase 2: Storage layer widening
Widen internal storage layer types from int to long for segment sizes and physical file positions:
FileRecords: Internal AtomicInteger changed to AtomicLong for size tracking. New sizeInBytesLong(), sliceLong(), truncateToLong() methods added for callers that need long precision. The existing BaseRecords.sizeInBytes() interface remains int to avoid cascading changes.
LogSegment: New sizeInBytesLong() alongside existing size(). Methods recover(), append(), shouldRoll(), read(), and truncateTo() widened to use long for positions and sizes.
OffsetIndex dual format: The OffsetIndex supports two entry formats, controlled by a useLargeFormat constructor parameter:
- Legacy format (default): 8 bytes per entry (4-byte relative offset + 4-byte physical position). Used before MetadataVersion finalization.
- Large format: 12 bytes per entry (4-byte relative offset + 8-byte physical position). Used after MetadataVersion finalization. Supports segment positions greater than 2 GB.
The default is legacy format (useLargeFormat=false). All production code paths (via LazyIndex, RemoteIndexCache, DumpLogSegments) create OffsetIndex instances in legacy format. The large format is only used when explicitly opted in via the 5-arg constructor after MetadataVersion verification.
The AbstractIndex base class is enhanced with an effectiveEntrySize() method that safely resolves the entry size at construction time without calling overridable methods from the constructor.
MetadataVersion gating: A new MetadataVersion entry IBP_4_4_IV1(32, "4.4", "IV1", true) gates the format change. The didMetadataChange=true flag ensures downgrade is blocked after finalization. The helper method isLargeIndexFormatSupported() returns true when the cluster MetadataVersion is at or above IBP_4_4_IV1.
RemoteStorageManager API: New default methods added to the RemoteStorageManager interface with long position parameters. These delegate to the existing int methods with bounds checking. Existing RemoteStorageManager implementations continue to work unchanged. The old int methods are marked @Deprecated.
Cascading type changes: Other code that consumes sizeInBytes() or OffsetPosition.position is updated. Key call sites include UnifiedLog, Cleaner, LogLoader, LocalLog, RemoteLogManager, RemoteIndexCache, DelayedFetch, and DumpLogSegments.
Compatibility, Deprecation, and Migration Plan
Rolling upgrade path
- Upgrade all brokers to the new version. Do not finalize the metadata version yet. Brokers write index files in legacy 8-byte format (identical to old brokers). Full backward and forward compatibility. Downgrade is safe.
- Finalize the metadata version via
kafka-features.sh upgrade --release-version 4.4. New index files are written in 12-byte format. Existing 8-byte index files continue to be read correctly. When an index is rebuilt (for example duringLogSegment.recover()), it is written in the new format. Thesegment.bytesconfig upper bound is lifted. - Downgrade after finalization is blocked because
IBP_4_4_IV1hasdidMetadataChange=true, consistent with existing KRaft downgrade rules.
Backward compatibility
- Config parsing:
INTvalues stored as strings (for example"1073741824") parse correctly asLONG. No user action required. - Index format: Before MetadataVersion finalization, all indexes use the legacy 8-byte format. Old and new brokers produce identical index files.
- RemoteStorageManager: Existing implementations only implement the
intmethods. The newlongdefault methods delegate to the oldintmethods with bounds checking. No changes required for existing RSM plugins.
Forward compatibility
- Before finalization: full downgrade is safe. All index files are in legacy format.
- After finalization: downgrade is blocked by KRaft metadata version rules.
- If segments larger than 2 GB were created after finalization, operators must set
segment.bytesback to 2 GB or less and wait for segment rolls before downgrading.
Deprecation
RemoteStorageManager.fetchLogSegment(RemoteLogSegmentMetadata, int)is deprecated in favor offetchLogSegment(RemoteLogSegmentMetadata, long).RemoteStorageManager.fetchLogSegment(RemoteLogSegmentMetadata, int, int)is deprecated in favor offetchLogSegment(RemoteLogSegmentMetadata, long, long).- No configs are deprecated or removed.
- The deprecated methods will be removed in a future major release.
Impact on existing users
- Users who never set
segment.bytesabove the current default (1 GB) are completely unaffected. - Users who want larger segments must first finalize
MetadataVersiontoIBP_4_4_IV1. - Custom
RemoteStorageManagerimplementations should migrate to thelongoverloads at their convenience. The deprecatedintmethods continue to work.
Test Plan
Unit tests
- Config parsing and round-trip tests verify
LONGtype works correctly forlog.segment.bytesandsegment.bytes. - OffsetIndex dual-format tests verify both 8-byte legacy and 12-byte large entry formats can be written and read correctly.
- OffsetIndex upgrade/downgrade tests verify that legacy-format indexes can be read by new code (before MetadataVersion finalization), and that indexes written by new code in legacy mode can be read by old code.
- OffsetIndex migration test verifies the full path: legacy format, reset, rebuild with positions greater than
Integer.MAX_VALUE. - FileRecords tests verify
sizeInBytesLong()returns accurate values exceedingInteger.MAX_VALUE, andtruncateToLong()handles truncation amounts exceedingInteger.MAX_VALUE. - LogSegment tests verify
shouldRoll()works withmaxSegmentBytesgreater than 2 GB,sizeInBytesLong()consistency withsize()for small segments, and recovery preserves all records. - RemoteLogManager tests verify the new
long-paramfetchLogSegment()methods work correctly.
Integration tests
- Tiered storage integration tests (OffloadAndConsumeFromLeaderTest, RollAndOffloadActiveSegmentTest, DynamicSegmentSizeChangeTest) verify end-to-end produce, offload, consume, and broker bounce with mixed segment sizes.
- DynamicSegmentSizeChangeTest specifically tests dynamically changing segment size on a tiered-storage-enabled topic and verifying data integrity after broker restart.
System-level verification (performed manually)
- Live end-to-end test: 300 MB/s produce with dynamic config changes (1 GB to 4 GB to 1 GB segment sizes). Verified 4 GB segments are created, consumer reads across mixed segment sizes, and data integrity is maintained.
- Upgrade test: Old broker (trunk) produces 3000 records across 3 segments. New broker starts on same data directory, reads old data, produces new records, all produce and consume operations succeed. Index files remain in legacy 8-byte format.
- Downgrade test: After upgrade test, old broker starts on data written by new broker. Reads all data (both pre-upgrade and post-upgrade), produces new records, all operations succeed.
Rejected Alternatives
1. Keep log.segment.bytes as INT permanently
Rejected. The 2 GB limit is an artificial constraint from a type choice made when storage hardware was smaller. Modern deployments routinely manage multi-TB partitions where 2 GB segments create excessive overhead in file handles, segment rolls, compaction cycles, and tiered storage operations.
2. Add a separate log.segment.bytes.long config
Rejected. Maintaining two configs for the same purpose adds confusion for operators. A single config with a type change is cleaner and follows the precedent set by KIP-1161, which reclassified several configs from STRING to LIST type.
3. Widen index entries to 16 bytes (8-byte offset + 8-byte position)
Rejected. The relative offset (4 bytes) is sufficient because it represents the delta from the segment base offset, not an absolute offset. Only the physical position needs widening to 8 bytes. Using 16 bytes per entry would waste 50% more space for no practical benefit.
4. Use unsigned int for physical position (4 GB range)
Rejected. Java does not natively support unsigned integers, making the code error-prone (values above Integer.MAX_VALUE appear negative, breaking comparison operators and binary search). The additional 2 GB headroom is not worth the complexity. Widening to long is the clean solution and future-proofs the format.
5. Do only the config change without the index format change
Rejected as the complete approach. While Phase 1 (config type change with INT range cap) is useful as a stepping stone, it does not deliver the actual user-facing value of larger segments. A single KIP covering the full scope ensures the community reviews the complete design, even though implementation can be phased across multiple PRs.
6. Per-partition marker file for index format migration
Rejected. An earlier prototype used a .index_version marker file in each partition directory to track whether indexes had been rebuilt in the new format. On first startup after upgrade, if the marker was absent, all indexes were rebuilt. This approach was rejected because:
- It does not follow Kafka conventions. Every other on-disk format change in Kafka uses MetadataVersion gating.
- No downgrade support. The marker file approach immediately writes new-format indexes on upgrade, with no way to revert. MetadataVersion gating allows all brokers to be upgraded (still writing old format) before the format switch is finalized.
- Mixed-version cluster risk. In a rolling upgrade, broker A would immediately start writing 12-byte indexes while broker B (not yet upgraded) still expects 8-byte.
- False negatives in format detection. Old 8-byte index files whose size is divisible by both 8 and 12 (for example, 72 bytes from 9 entries) cannot be reliably detected by file-size heuristics alone, leading to silent data corruption on approximately 33% of index files during upgrade.
7. Magic byte header in index files for format detection
Rejected. Adding a version byte at the start of each index file would allow self-describing format detection, but it adds complexity to the read path (must check header before every index open), and old brokers would misread the header byte as part of the first entry, potentially producing garbled offset lookups before sanity checks catch it. MetadataVersion gating is simpler and avoids these edge cases entirely.