Current state: Under Discussion
Discussion thread: https://lists.apache.org/thread/vdp8scrrzdq7ofvl0mm84dhphq8kmzgc
Vote thread:https://lists.apache.org/thread/5280h24g205vn69dxr44lc15dt3ncrvz
JIRA:
In production, producers commonly send keyed messages to leverage semantic partitioning for ordering guarantees, stream joins, and cross-cluster replication. However, kafka-producer-perf-test always produces records with null keys, making benchmark results systematically optimistic and unable to reflect the performance characteristics of real keyed workloads.
This proposal adds key distribution support to kafka-producer-perf-test, allowing engineers to benchmark keyed workloads with configurable key ranges and distribution strategies.
This proposal adds two new command-line arguments to kafka-producer-perf-test:
Controls how message keys are assigned:
Defines the size of the key space. Must be a positive integer.
public enum KeyDistribution {
NONE, RANGE, RANDOM
} |
Keys are serialized as their decimal string representation encoded in UTF-8, consistent with the ByteArraySerializer already configured for the producer. This keeps keys human-readable in tools like kafka-console-consumer.
| Distribution | Key value |
|---|---|
| NONE | null |
| RANGE | Integer.toString(recordIndex % keyRange) |
RANDOM | Integer.toString(random.nextInt(keyRange)) |
Performance note: The random distribution reuses a single SplittableRandom instance that is already constructed for payload generation. SplittableRandom.nextInt() is a lightweight, non-thread-safe PRNG with no allocation overhead, so key generation adds negligible latency to the hot path.
ConfigPostProcessor enforces mutual consistency between the two new arguments:
| Condition | Error |
|---|---|
| --key-distribution range or random without --message-key-range | --message-key-range is required when --key-distribution is 'range' or 'random'. |
| --message-key-range specified with --key-distribution none | --key-distribution must be 'range' or 'random' when --message-key-range is specified. |
| --message-key-range ≤ 0 | --message-key-range should be greater than zero. |
bin/kafka-producer-perf-test.sh \ --topic my-topic --num-records 1000000 --record-size 1024 \ --throughput -1 --bootstrap-server localhost:9092 |
bin/kafka-producer-perf-test.sh \ --topic my-topic --num-records 1000000 --record-size 1024 \ --throughput -1 --bootstrap-server localhost:9092 \ --key-distribution range --message-key-range 100 |
bin/kafka-producer-perf-test.sh \ --topic my-topic --num-records 1000000 --record-size 1024 \ --throughput -1 --bootstrap-server localhost:9092 \ --key-distribution random --message-key-range 10000 |
The default value of --key-distribution is none, which preserves the current behavior of sending null-key records. Existing scripts and benchmarks continue to work without modification.
All remaining tests should pass, and new unit test.
An alternative design would use UUID.randomUUID().toString() as the key for the random distribution, providing globally unique keys with no repeated values across the entire benchmark run.
This was rejected for two reasons:
UUID.randomUUID() uses SecureRandom internally, which is significantly slower than SplittableRandom.nextInt() and could become a bottleneck in high-throughput benchmarks — the opposite of what a perf tool should do.Engineers who genuinely need globally unique keys can use --key-distribution random --message-key-range <large-number> (e.g., 2^31−1) to approximate the same effect without the overhead.