A producer ID (PID) is a 64-bit identifier assigned by the broker to each idempotent or transactional producer. It has three key properties:
NO_PRODUCER_ID) marks non-idempotent batches.Without PID mapping, two independent clusters can assign the same producer ID to different producers. When records from both source clusters are mirrored into the same destination partition, the ProducerStateManager (PSM) sees two unrelated producers sharing one PID.
The following PID transformation is applied before appending mirrored data to the destination cluster:
-(PID + 2) |
This is a simple negation that maps all non-negative PIDs into the negative space. The +2 offset avoids mapping PID 0 to 0 and keeps PID -1 (non-idempotent) untouched.
There are a couple of problems with this approach that are evident when looking at the chained mirroring use case.
When B mirrors to C, PIDs already negative from A get re-transformed: -((-7) + 2) = 5 , which restores the original PID and collides with local producers on C.
A B C D
-1 -------> -1 -------> -1 -------> -1
5 -------> -7 --------> 5 -------> -7
5 -------> -7 # collision |
Even if we make the mapping idempotent by skipping negative PIDs, when A has local PID 5 and B also has local PID 5, both map to -7 on any downstream cluster. These are different producers, but they become indistinguishable. The PSM cache stores the transformed PID with no awareness of its origin, so a collision silently overwrites the previous entry breaking txn consistency within the log.
A B C 5 -------> -7 -------> -7 5 -------> -7 # collision |
We identified the following scenarios caused by interleaving records from a PID collision:
OutOfOrderSequenceException. Batches are silently accepted under the same PID as they are coming from the leader (append origin == REPLICATION). The PSM cache entry is updated with whatever arrives last. Silent data corruption with zero signals, not even a warning.First of all the PID mapping logic is updated to skip it for positive PIDs, making it idempotent. Rather than addressing collisions at runtime, this approach proactively eliminates stale state on failover. When mirroring stops, no mirrored producer remains active: the fetcher has been removed and the partition is about to become writable. All negative PIDs in the ProducerStateManager (PSM) are stale and can be safely expired. A MIRROR_PID_RESET control record (type 7) is written to each destination partition's log during the STOPPING state transition, after the fetcher has been removed and truncation to last stable offset is complete. The dump tool is enhanced to deserialize and display MIRROR_PID_RESET records. Control batches are filtered out by the consumer fetcher, so the barrier is invisible to application consumers, just like transaction markers.
The key is a standard control record key (version=0, type=7), while the value has the following schema:
{
"type": "data",
"name": "MirrorPidResetRecord",
"validVersions": "0",
"flexibleVersions": "0+",
"fields": [
{ "name": "Version", "type": "int16", "versions": "0",
"about": "The version of the mirror PID reset record."},
{ "name": "SourceClusterId", "type": "string", "versions": "0",
"about": "The source cluster UUID for verification."}
]
} |
When the leader writes this barrier:
On replica recovery or log loading, the barrier is replayed from the log, triggering the same PSM expiration on followers. This ensures all replicas converge to the same clean producer state.
With this setup, all negative PIDs are always coming from a single source cluster where no collision is possible.
When we have a mirroring chain from A to B to C :
After a failover from A to B, the operator may later reverse the direction and mirror B back to A (failback).
When A begins mirroring from B: