Versions Compared

Key

  • This line was added.
  • This line was removed.
  • Formatting was changed.

...

Rather than chasing down problems in our mainline branches after the fact, or requiring up-to-date PRs in GitHub, this KIP proposes to leverage the Merge Queue feature of GitHub.

The merge queue effectively gives us the "require up-to-date" feature without the need for constant rebuildingacts as a gatekeeper for trunk. It does so by placing all PRs into a queue rather than merging them directly. Once in the queue, a Continuous Integration (CI) job will run that applies the PR on top of the base branch and then runs our a custom job. If successful, the PR is merged to the base branch and the next PR in the queue is built. If the our custom job is not successful, the PR remains open and the author is automatically notified of the failure. This process continues as long as there are pending PRs to be merged.

The custom job acts as a gatekeeper for trunk. We will start with something similar to our existing "Compile and Check Java" workflowThis controlled and sequential workflow will allow us to ensure that no change made to Kafka will break our mainline branches.

GitHub Details

When enabled, the merge queue replaces the "Squash and Merge" button that we use today. We can still do a squash merge, but the merge button becomes "Merge when ready". Clicking this button is tantamount to merging the PR to trunk. Similar to today, only committers are authorized to click the button, and only PRs which have been approved should be merged.

...

Figure 1: The GitHub UI for merging a Pull Request with the Merge Queue

...




There are a few important configurations of GitHub's Merge Queue that we must understand.

...

As Pull Requests (PRs) are enqueued, they can run the validation CI check concurrently to avoid blocking. These concurrent builds are still serialized in the sense that a later PR in the queue will include the earlier PRs. Setting "Build concurrency" higher than 1 will allow multiple CI runs to happen concurrently.

...

The extent of validation we can perform in the merge queue job is governed by the rate of PRs we expect to merge. Taking data from https://github.com/apache/kafka/graphs/commit-activity, we can see that over the last year we have merged around 15 commits per day. Our "Compile and Check Java" CI step is very consistently taking 11-12 minutes on trunk (with no caching). Our full test suite runs in around 2 hours on trunk (again, no caching). With the current rate of change in Kafka, we cannot reasonably run the full test suite for each PR in the merge queue. However, it would quite reasonable to perform compilation and static checks on each PR.

Since there are many unknowns with the merge queue, this KIP proposes that we enable it in a simple configuration with our "Compile and Check Java" step. This will give us the most immediate benefit of the merge queue which is protecting trunk from being brokenwithout overly complicating things. By disabling batching and concurrency initially, we can simplify the mental model of this new change management. This will give the community a chance to adapt to this new paradigm and really see how it works. 

...