You are viewing an old version of this page. View the current version.

Compare with Current View Page History

Version 1 Next »

This page is meant as a template for writing a KIP. To create a KIP choose Tools->Copy on this page and modify with your content and replace the heading with the next KIP number and a description of your issue. Replace anything in italics with your own description.

Status

Current state: DRAFT

Discussion thread: here [Change the link from the KIP proposal email archive to your own email thread]

JIRA: here

Please keep the discussion on the mailing list rather than commenting on the wiki (wiki discussions get unwieldy fast).

Motivation

Often users want to reduce the cardinality of a stream (aggregate) independent of time semantics so that downstream topologies are not overwhelmed by the throughput of upstream producers. A classic example is to aggregate together N seconds of events for a single key into a single one so that a downstream consumer that does something heavy per-event is not overloaded (e.g. it makes a remote call):

stream
	.windowBy(TimeWindows.ofSizeAndNoGrace(Duration.ofSeconds(10))
	.aggregate((k, v, kvList) -> kvList.add((k, v)))
	.suppress(Suppressed.untilWindowCloses(...))

This approach, however, makes handling late arriving data intractable. The user is forced between two choices:

  1. Use a tumbling window with grace period of 0 and ignore out-of-order events
  2. Use a tumbling window with a grace period large enough to account for late arriving data and suppress for much longer than desired, introducing excessive end-to-end latency (and potentially undue memory pressure)

Typically the behavior that users want with this type of operation is to group keys together until a certain amount of time passes and emit them as a single record, always accepting late arriving records and just putting them into whatever the “current” window is. Today the only way to implement this semantic is using a punctuation and manually storing events in an aggregation (which is inefficient for many reasons and can cause severe bottlenecks).

Public Interfaces

Briefly list any new interfaces that will be introduced as part of this proposal or any existing interfaces that will be removed or changed. The purpose of this section is to concisely call out the public contract that will come along with this feature.

We will be introducing a new type of Windows that always computes the current window based on stream time instead of event time.

public final class StreamTimeWindows extends Windows<TimeWindow> {
		
		/**
		 * Return a window definition with the given window size.
		 * <p> 
		 * This represents stream time window semantics, which are fixed-size,
		 * gap-less, non-overlapping windows. Static windows use the current
		 * stream time when determining which window an event fits into instead
		 * of the event time (as the other window types use).
		 * <p>
		 * Note that stream time windows will always accept late arriving data, and
		 * should only be used in situations where the ordering of data is not
		 * necessarily important to the semantics of your application.
		 */
		public static StreamTimeWindows ofSize(final Duration size) { ... }
    
}

Here is an example of how events would map to windows:

The example in the motivation section then becomes:

stream
	.windowBy(StreamTimeWindows.ofSize(Duration.ofSeconds(10))
	.aggregate((k, v, kvList) -> kvList.add((k, v)))
	.suppress(Suppressed.untilWindowCloses(...))

And all events that come within 10 seconds of stream time (independent of their event time) will be aggregated together, suppressed and emitted when the stream time elapses by 10 seconds.

Proposed Changes

Describe the new thing you want to do in appropriate detail. This may be fairly extensive and have large subsections of its own. Or it may be a few sentences. Use judgement based on the scope of the change.

First, we need to expand the StreamTimeWindows class to take in the current stream time when computing the window for a given event:

public abstract class Windows<W extends Window> {
  ...
  public abstract Map<Long, W> windowsFor(
    final long timestamp, 
    final long observedStreamTime
  );
}

The implementation of StreamTimeWindows will then use the observedStreamTime when computing which window the current event should fall into:

@Override
public Map<Long, TimeWindow> windowsFor(
  final long timestamp, 
  final long observedStreamTime
) {
	long windowStart = (observedStreamTime / sizeMs) * sizeMs;
  return Map.of(windowStart, new TimeWindow(windowStart, windowStart + sizeMs));
}

The gracePeriodMs for StreamTimeWindows is always zero, which means there will only ever be one open window per key and events will always fall within this window.

Compatibility, Deprecation, and Migration Plan

N/A

Test Plan

Describe in few sentences how the KIP will be tested. We are mostly interested in system tests (since unit-tests are specific to implementation details). How will we know that the implementation works as expected? How will we know nothing broke?

Nothing special needs to be tested beyond unit and integration tests to make sure we cover situations like out of order and late arriving events.

Rejected Alternatives

If there are alternative ways of accomplishing the same thing, what were they? The purpose of this section is to motivate why the design is the way it is and not some other way.

  • Instead of introducing a new window type, we could add a property like Strictness  on the existing TimeWindows  type. This would end up being more confusing than the proposed API since grace period do not apply to the type of windows suggested here.
  • No labels