Skip to content

Skip graceful worker retirement during cluster teardown - #9368

Open
abhisek343 wants to merge 2 commits into
dask:mainfrom
abhisek343:fix/skip-retirement-on-cluster-close
Open

abhisek343 wants to merge 2 commits into
dask:mainfrom
abhisek343:fix/skip-retirement-on-cluster-close

Conversation

@abhisek343

Copy link
Copy Markdown

Closes #9300

Full "SpecCluster" teardown currently calls "scheduler_comm.retire_workers()" before closing workers, even after the cluster has entered "Status.closing". This can trigger unnecessary data replication/drop work immediately before all workers are removed.

This change skips scheduler retirement only during full cluster teardown. Normal scale-down continues to retire workers gracefully.

A regression test verifies that closing the cluster:

  • does not call "retire_workers()"

  • still closes the worker

  • removes the worker from cluster state

  • Tests added

  • Tests passed

  • Passes "pixi run lint"

@github-actions

Copy link
Copy Markdown
Contributor

Unit Test Results

See test report for an extended history of previous test failures. This is useful for diagnosing flaky tests.

    40 files  ± 0      40 suites  ±0   14h 34m 40s ⏱️ + 4m 31s
 4 161 tests + 1   3 981 ✅ ±0    178 💤 ±0   2 ❌ + 1 
80 979 runs  +18  76 719 ✅ ±0  4 239 💤  - 1  21 ❌ +19 

For more details on these failures, see this check.

Results for commit c60b590. ± Comparison against base commit dc182bd.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SpecCluster._correct_state_internal closes workers with Nanny's default 5s timeout; shutdown often fails on LocalCluster(processes=True)

1 participant