Visitar URL original
[Bug report] Cancelling a local job can hold a worker thread forever or report SUCCEEDED · Issue #13678 · apache/gravitino · GitHub
Skip to content

[Bug report] Cancelling a local job can hold a worker thread forever or report SUCCEEDED #13678

Description

@LuciferYang

Version

main branch

Describe what's wrong

LocalJobExecutor.cancelJob stops a running job with Process.destroy(), which sends SIGTERM, and then marks the job CANCELLING. Two things can go wrong after that.

A job whose process ignores SIGTERM never exits. runJob keeps blocking in process.waitFor(), the job stays CANCELLING, and its worker thread is never released. The worker pool is fixed at maxRunningJobs, so each such job takes a slot for good; once every slot is held this way, new jobs stay queued until the server restarts.

A job whose process traps SIGTERM and exits 0 is reported SUCCEEDED. runJob checks the exit code before it checks whether the job was being cancelled, so a clean exit taken in response to the cancel wins over the cancel.

Error message and/or stacktrace

No error is raised. In the first case the job stays CANCELLING indefinitely; in the second it ends SUCCEEDED instead of CANCELLED.

How to reproduce

  1. Submit a shell job running trap '' TERM; while :; do sleep 1; done, wait until it is STARTED, and cancel it. It never leaves CANCELLING.
  2. Submit a shell job running trap 'exit 0' TERM; while :; do sleep 1; done, wait until it is STARTED, and cancel it. It ends SUCCEEDED.

Activity

added a commit that references this issue on Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

2.0.0Release v2.0.0

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions