DUE TO SPAM, SIGN-UP IS DISABLED. Goto Selfserve wiki signup and request an account.
...
- Unique: uniquely identifies a producer for idempotent deduplication and transaction tracking.
- Stable: once assigned, a PID persists across producer sessions (for transactional producers) or until expiration.
- Non-negative: valid PIDs are >= 0. The value -1 (NO_PRODUCER_ID) marks non-idempotent batches.
Without PID mapping, two independent clusters can assign the same producer ID to different producers. When records from both source clusters are mirrored into the same destination partition, the ProducerStateManager (PSM) sees two unrelated producers sharing one PID.
...
We identified the following scenarios issues caused by interleaving records from a PID collision:
- Same epoch, wrong sequence: No OutOfOrderSequenceException. Batches are silently accepted under the same PID as they are coming from the leader (append origin == REPLICATION). The PSM cache entry is updated with whatever arrives last. Silent data corruption with zero signals, not even a warning.
- Different epochs: No fencing exception. Lower epoch batch is accepted with a warning log. Both producers coexist under the same PID. Silent corruption, only a WARN log line as a hint.
- Transactional interleaving: Commit/abort markers from one producer close the other's transaction. No exception. Silent transaction corruption.
...
Rather than transforming PIDs at write time, this approach proactively expires stale producer state on failover. The key insight is that during mirroring, the destination partition is read-only: no local producers exist, so all PSM entries originate from mirrored data. When mirroring stops, all PSM entries are stale and can be safely expired. Records from the source are stored as-is on the destination, with no PID modification, which otherwise would require a checksum recalculation. A MIRROR_PID_RESET control RESET control batch (type 7) is written to each destination partition's log during the STOPPING state transition, after the fetcher has been removed and truncation to last stable offset is complete, but before the partition becomes writable.
| Code Block |
|---|
A B C
5(A) -------> 5(A) -------> 5(A) # 5(A) means PID:5, source cluster:A
CB ---------> CB # Control Batch appended when A failover to B
5(B) -------> 5(B) # In B and C, even if 2 records with PID 5, they won't duplicate with each other because of the control batch. |
The mirror partition state transitions are:
- STOPPING: remove fetchers, truncate to LSO, persist last mirrored offsets, write MIRROR_PID_RESET barrier
- STOPPED: partition is writable (terminal state, no actions)
The key follows the standard control record format (version=0, type=7). The value uses the MirrorPidResetRecord schemathe MirrorPidResetRecord schema:
| Code Block |
|---|
{
"type": "data",
"name": "MirrorPidResetRecord",
"validVersions": "0",
"flexibleVersions": "0+",
"fields": [
{ "name": "Version", "type": "int16", "versions": "0",
"about": "The version of the mirror PID reset record."},
{ "name": "SourceClusterId", "type": "string", "versions": "0",
"about": "The source cluster UUID for verification."}
]
} |
The SourceClusterId field SourceClusterId field records which source cluster the mirrored data came from, enabling future validation (e.g. detecting unexpected source cluster changes) and data provenance tracing from the log itself.
When the barrier batch is encountered during append or during log recovery, all producer entries are removed from the PSM. This ensures leaders, followers, and recovery all handle the barrier consistently. Because the partition is read-only during mirroring, all PSM entries originate from mirrored data. Expiring all entries is safe: no local producer state exists to preserve. Control batches are filtered out by the consumer fetcher via isControlBatch via isControlBatch checks. The barrier is invisible to application consumers, just like transaction markers (commit/abort). The log dump tool is enhanced to deserialize and display MIRROR_PID_RESET records.
Chaining example
...
records
...
.
...
The barrier approach works correctly with all practical mirroring topologies:
...