DUE TO SPAM, SIGN-UP IS DISABLED. Goto Selfserve wiki signup and request an account.
Status
Current state: Under Discussion
Discussion thread: https://lists.apache.org/thread/t7fbxcr41lxojydrsvy569bgfm91o3j6
JIRA:
KAFKA-7699
-
Getting issue details...
STATUS
Please keep the discussion on the mailing list rather than commenting on the wiki (wiki discussions get unwieldy fast).
Motivation
Kafka Streams do not provide a way to easily trigger periodic callbacks at specific times. Wall-clock time punctuation allows to schedule periodic callbacks based on wall-clock time progress, but the punctuation time starts when the punctuation is scheduled. As a result, the callback is triggered at a non-deterministic time. It would be nice to allow a punctuation to be triggered at a fixed/anchored time, independent of when the punctuation was registered. For instance, this will allow for triggering a punctuation at the start of every hour, i.e. HH:00:00.
Real life use cases can be taken from the energy sector. The automatic closed loop balancing of the power grid requires calculations to be published at the start of every 10th second. Today, this is solved by a work-around utilising a "has published this interval"-flag, and a state store to ensure a punctuation is only triggered once per interval:
class PunctuateProcessor(
private val hasForwardedStoreName: String,
private val forwardTime: Duration,
) : ContextualProcessor<String, avro_value, String, avro_value>(),
ILogging by Logging<PunctuateProcessor>() {
private lateinit var hasForwardStore: WindowStore<String, Boolean>
private lateinit var forwardSchedule: Cancellable
private val hasForwardStoreKey = "hasPublished"
override fun init(context: ProcessorContext<String, avro_value>) {
super.init(context)
this.hasForwardStore = context.getStateStore(hasForwardedStoreName)
forwardSchedule =
context().schedule(Duration.ofMillis(500), PunctuationType.WALL_CLOCK_TIME) { forwardRecordsIfTime() }
}
override fun process(record: Record<String, avro_value>) {
// Store incoming records
}
private fun forwardRecordsIfTime() {
val currentTime = Duration.ofMillis(context().currentSystemTimeMs())
val flooredTime = Duration.ofMillis(floorTo10Second(currentTime.toMillis()))
if (isTimeToForward(currentTime) && !hasForwardedThisInterval(flooredTime)) {
forwardRecords()
hasForwardStore.put(hasForwardStoreKey, true, flooredTime.toMillis())
}
}
private fun isTimeToForward(currentTime: Duration): Boolean = (currentTime.toSecondsPart() % 10) >= (forwardTime.toSecondsPart() % 10)
private fun hasForwardedThisInterval(intervalStart: Duration): Boolean =
hasForwardStore.fetch(hasForwardStoreKey, intervalStart.toMillis()) ?: false
Generally, the work-around can be characterised as flag-polling, where the poll frequency is determined by a wall-clock punctuation. Hence, the punctuation trigger time is only as precise as the wall-clock punctuation trigger interval. A more frequent wall-clock punctuation will give a more precise trigger time, at the cost of firing an increased number of flag checks (i.e. punctuations). This is a trade-off that can be removed with anchored punctuations.
Proposed Changes
The goal is to introduce an anchored wall-clock punctuation that will have functionally similarities with that of running a cron job. The proposed API change is to extend the `schedule()` method in the `ProcessingContext` interface with a parameter to represent the start time in epoch milliseconds for the schedule's anchored time, without enforcing any changes upon the existing users:
package org.apache.kafka.streams.processor.api;
public interface ProcessingContext {
// New method allowing for anchored punctuation
Cancellable schedule(final Duration interval, final long startTime, final PunctuationType type, final Punctuator callback);
// Existing method
Cancellable schedule(final Duration interval, final PunctuationType type, final Punctuator callback) {
schedule(interval, null, type, callback);
}
}
The `startTime`, together with the `interval` and the wall clock, will determine the next trigger time for the callback. Hence, the anchored punctuation is only supporting wall clock (i.e. stream time) in this first iteration.
The method for calculating the next, anchored trigger time, could look something like:
long currentTime = currentSystemTimeMillis();
// If currentTime is before startTime, return the difference between startTime and currentTime
if (currentTime < startTime) {
return startTime - currentTime;
}
// Calculate how many intervals have passed since startTime
long elapsedTime = currentTime - startTime;
// Calculate how many full intervals have passed
long intervalsPassed = elapsedTime / interval;
// Calculate the time of the next trigger time
long nextTriggerTime = startTime + (intervalsPassed + 1) * interval;
// If the current time is already a trigger time, return 0
return nextTriggerTime - currentTime;
Compatibility, Deprecation, and Migration Plan
The goal is to not enforce any changes upon the existing users of the schedule method, but instead expand the current schedule options with a new anchored schedule option. This can be achieved by using method overloading.
Test Plan
The plan is to test the new schedule option in the same way that the current schedule options are tested. The anchored wall-clock punctuation is a new feature, and the feature should therefore not affect any current features or users.
Rejected Alternatives
Cron job
Creating a new schedule method that takes in a cron job expression as a parameter:
package org.apache.kafka.streams.processor.api;
public interface ProcessingContext {
// New method allowing for anchored punctuation using cron expressions
Cancellable schedule(final String cronExpression, final Punctuator callback);
// Existing method
Cancellable schedule(final Duration interval, final long startTime, final PunctuationType type, final Punctuator callback);
}
The usage of cron expressions would require the inclusion of a new dependency, such as Quartz. Generally, we wish to avoid bringing in new dependencies if it is avoidable. Also, it is wise to start simple when implementing a new future, and then build up the feature incrementally.