The continuous integration build is critical to the project's health and development velocity. It enables contributors to make changes to Spark with the confidence that they have not inadvertently broken anything.
There are several considerations to keep in mind when making changes to the build.
The Apache Software Foundation (ASF) imposes resource usage limits on GitHub Actions, including on the number of concurrent running jobs as well as on the total number of run minutes used per week and rolling five-day window. When we hit these limits it can cause jobs to queue up, bringing development to a crawl.
Workflows triggered by on: pull_request count against the base repo's (i.e. this repo's) activity limits. For this reason it's important to use on: push for heavy workflows. on: push workflows count against the fork's limits and not ours, enabling more jobs to run concurrently.
The pull_request workflow trigger provides key benefits that make working with workflows easier, including:
- The workflow runs against the target repo's context (i.e.
apache/spark). - GitHub provides a stable merge reference.
- PR metadata is readily available via
github.event.pull_request.
Without these benefits, some GitHub Actions patterns become more difficult or impossible to implement. We've implemented our own custom checkout pattern via the checkout-and-sync composite action.
If in the future GitHub somehow allows workflows to run on: pull_request while counting those jobs against the forks' limits instead of ours, that would help eliminate various distortions of our build caused by the move away from on: pull_request. We would get the best functionality from GitHub Actions without hitting job run limits.
If you find yourself repeating the same set of instructions across multiple jobs, extract them into a reusable composite action under actions/.
Keep these security considerations in mind when authoring workflows, as it's possible for malicious forks to use poorly written workflows to steal secrets or merge unapproved code.