DUE TO SPAM, SIGN-UP IS DISABLED. Goto Selfserve wiki signup and request an account.
...
Generally, the work-around can be characterised as flag-polling, where the poll frequency is determined by a wall-clock punctuation. Hence, the punctuation trigger time is only as precise as the wall-clock punctuation trigger interval. A more frequent wall-clock punctuation will give a more precise trigger time, at the cost of firing an increased number of flag checks (i.e. punctuations). This is a trade-off that can be removed with anchored punctuations.
...
Public API Interfaces
The goal is to introduce an anchored wall-clock punctuation that will have functionally similarities with that of running a cron job. In short, the user should be allowed to specify the start time for the schedule, instead of relying on the non-deterministic time for when the schedule was registered.
...
| Code Block | ||
|---|---|---|
| ||
package org.apache.kafka.streams.processor.api;
public interface ProcessingContext {
// New method allowing for anchored punctuation
Cancellable schedule(final Duration interval, final longInstant startTime, final PunctuationType type, final Punctuator callback);
// Existing method
Cancellable schedule(final Duration interval, final PunctuationType type, final Punctuator callback) {
schedule(interval, null, type, callback);
}
} |
Proposed changes
PunctuationType support
It is planned to support both system time and stream time in this first iteration, as it is deemed more of a hassle to gracefully handle non-supported punctuation types rather than implementing support for both punctuation types.
Punctuation semantics
The startTimeThe `startTime`, together with the `interval` interval and the wall clockcurrent time, will determine the next trigger time for the callback. It is planned that only wall clock (i.e. stream time) is supported in this first iteration. The method for calculating the next , anchored trigger time , could look something like:
| Code Block | ||||
|---|---|---|---|---|
| ||||
long currentTime = currentSystemTimeMillis(); // or currentStreamTime(); long startTimeMillis = startTime.toEpochMillis(); if (currentTime < startTimestartTimeMillis) { return startTimestartTimeMillis - currentTime; } long elapsedTime = currentTime - startTimestartTimeMillis; long intervalsPassed = elapsedTime / interval; long nextTriggerTime = startTimestartTimeMillis + (intervalsPassed + 1) * interval; return nextTriggerTime - currentTime; |
Edge cases, and possibly unintuitive cases, will be described in the following paragraphs.
StartTime is defined in the past
If the startTime is defined as a point in time that is before the current time, the implementation will just skip forward and calculate the next trigger time.
If we have startTime at t=90, the current time is t=101, and the interval is 10s, the next trigger time will be t=110.
As a result, there will exist several schedule configurations yielding the same trigger times.
Starts and restarts
If the running Streams app start, or restarts, close to the scheduled trigger time, there is a possibility that callbacks will be missed.
Take for example a schedule configured with a start time of t=100 and an interval of 10 seconds. When the Streams app start up at t=101, it would wait until t=110 to trigger the first punctuation, and then proceed to punctuate as usual in 10s increments.
However, if the app happened to shut down at t=109, and did not come back up again before t=111, the app will miss the punctuation at t=110 and will have to wait until t=120 for the next punctuation.
Compatibility, Deprecation, and Migration Plan
...
The usage of cron expressions would require the inclusion of a new dependency, such as Quartz. Generally, we wish to avoid bringing in new dependencies if it is avoidable. Also, it is wise to start simple when implementing a new future, and then build up the feature incrementally. . The value-add of supporting cron expressions is also not clear, as schedules can be precisely configured "only" using a start time and an interval.