DUE TO SPAM, SIGN-UP IS DISABLED. Goto Selfserve wiki signup and request an account.
...
Current state: In Discussion
Discussion thread: here [Change the link from the KIP proposal email archive to your own email thread]
...
Kafka Stream supports different ways to window a stream for aggregation, but all of them are extremely sensitive to out of order timestamps and late arriving data. To handle this, we introduced the concept of a "grace period" by which a window can be retained for longer than the window end to allow for updating the result of a window with a late arriving recordrecord that arrives after the window end time has passed. This concept is powerful, but doesn't help users that want to aggregate based on offset order alone, not discarding any late arriving record.
...
- Use a tumbling window with grace period of
0and ignore out-of-order late arriving events - Use a tumbling window with a grace period large enough to account for late arriving data and suppress for much longer than desired, introducing excessive end-to-end latency (and potentially undue memory pressure)
...
We will be introducing a new type of Windows that always computes the current window based on the current stream time instead of event timestamp:
...
| Code Block |
|---|
public final class StreamTimeWindowsBatchWindows extends Windows<TimeWindow> { /** * Return a window definition with the given window size. * <p> * This represents stream time"batching" window semantics, which are fixed-size, * gap-less, non-overlapping windows. StaticBatched windows use the current * stream time when determining which window an event fits into instead * of the event time (as the other window types use). This allows users * to batch together records of the same key over a window of time. * <p> * The window boundaries are inclusive of the start point and exclusive, * of the end point. This means that a BatchWindows of size 1000 would have * windows from [0,1000),[1000,2000),etc... * <p> * Note that stream timebatch windows will always accept late arriving data, and * should only be used in situations where the ordering of data is not * necessarily important to the semantics of your application. */ public static StreamTimeWindowsBatchWindows ofSize(final Duration size) { ... } } |
Here is an example of how events would map to windows:
The example in the motivation section then becomes:
| Code Block |
|---|
stream .windowBy(StreamTimeWindowsBatchWindows.ofSize(Duration.ofSeconds(10)) .aggregate((k, v, kvList) -> kvList.add((k, v))) .suppress(Suppressed.untilWindowCloses(...)) |
...
First, we need to expand the StreamTimeWindows Windows class to take in the current stream time when computing the window for a given event:
| Code Block |
|---|
public abstract class Windows<W extends Window> {
...
public abstract Map<Long, W> windowsFor(
final long timestamp,
final long observedStreamTime
) {
return windowsFor(timestamp);
}
// as part of this KIP we will deprecate this method, which is technically public
// although users are not expected to implement this (it lives in a public package)
@Deprecated
public abstract Map<Long, W> windowsFor(final long timestamp);
} |
The implementation of StreamTimeWindows BatchWindows will then use the observedStreamTime when computing which window the current event should fall into:
| Code Block |
|---|
@Override
public Map<Long, TimeWindow> windowsFor(
final long timestamp,
final long observedStreamTime
) {
long windowStart = (observedStreamTime / sizeMs) * sizeMs;
return Map.of(windowStart, new TimeWindow(windowStart, windowStart + sizeMs));
} |
The gracePeriodMs for StreamTimeWindows BatchWindows is always zero, which means there will only ever be one open window per key and events will always fall within this window. Here is the documentation for the remaining public members of the class:
| Code Block |
|---|
/** * A fixed-size, stream-time based window specification used for aggregations. * <p> * The semantics of batched aggregations are: Every size() milliseconds, compute the aggregate total for the last * size() milliseconds. This is equivalent in semantics to TimeWindows with an advance of 0. * <p> * This class differs from TimeWindows in that the windows for a given event is computed based on the current observed * stream time as opposed to the timestamp of the given record. This effectively makes the windows ordered with respect * to the offset of the record instead of the timestamp and allows for batching together records within a window. */ public final class BatchWindows extends Windows<TimeWindow> { /** * Returns the window size, measured in milliseconds. */ public long size() { } /** * Returns 0, as offset ordered windows have no concept of "late arriving" records * because stream time is always monotonically increasing */ public long gracePeriodMs() { return 0L; } } |
Compatibility, Deprecation, and Migration Plan
...
- Instead of introducing a new window type, we could add a property like
Strictnesson the existingTimeWindowstype. This would end up being more confusing than the proposed API since grace period do not apply to the type of windows suggested here. - Names we discussed but didn't like:
GlobalWindow,StreamTimeWindow,ProcessTimeWindow, OffsetOrderedWindow. - Implementing a more general purpose batching mechanism that can batch based on number of records, wall clock or any other batching strategy. This is useful, but out of scope for this proposal as it introduces a different level of non-determinism that is not covered by the existing grace/windowing semantics. The way batching/suppression interrelate are also

