Repository navigation
launch: a Spot-reclaimed unit's further attempts go to RAPIDPIPE_BATCH_RECLAIM_QUEUE - #225
Merged
Merged
Conversation
…H_RECLAIM_QUEUE reconcile marks an attempt lost to a host reclaim (statusReason "Host EC2", job or last attempt) with scheduler_metadata.batch.reclaim. submit_unit then sends every further attempt of that unit, in the run or any run seeded from it, to RAPIDPIPE_BATCH_RECLAIM_QUEUE, so a reclaimed unit moves to on-demand capacity and never retries on Spot. Any other transient keeps the job queue. Unset, or equal to RAPIDPIPE_BATCH_JOB_QUEUE, nothing changes and no lookup is made. resolve_jobless searches both queues when they differ, counting distinct job ids. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
The launcher already records a Spot reclaim (FAILED, no container exit code,
statusReasonstartingHost EC2) astransientand resubmits the unit as a fresh attempt, but always on the process-wideRAPIDPIPE_BATCH_JOB_QUEUE, so a reclaimed unit could go straight back to Spot. This adds an optionalRAPIDPIPE_BATCH_RECLAIM_QUEUE:reconcilemarks an attempt lost to a host reclaim withscheduler_metadata.batch.reclaim = true(beside thejob_queueit already records). Other transients (exit 75,Cannot*,DockerTimeoutError*) are not marked.submit_unit(viaqueue_for_unit) sends every further attempt of a unit to the reclaim queue once any attempt of it, in the run or any run it was seeded from (units.seeded_from_unit, followed through every seed), was a reclaim. Sticky: a later exit 75 does not send it back to Spot. A never-reclaimed unit keeps the job queue.RAPIDPIPE_BATCH_JOB_QUEUE: the samesubmit_jobcall as before and no database lookup.run reconcile --resolve-joblesssearches both queues when they differ, counting distinct job ids (a queue's name and ARN are two spellings of one queue).Env contract:
RAPIDPIPE_BATCH_RECLAIM_QUEUE, when set and different fromRAPIDPIPE_BATCH_JOB_QUEUE, receives every further attempt of a unit reclaimed in its run or a run it was seeded from; unset or equal, behaviour is unchanged.Production stays unchanged: the processing-date loop (rapid_systems
op-processing-date-loop.sh) will export both variables asrapid-queue-prompt(on-demand) in a follow-up PR there. A future Spot test sets only the job queue torapid-queue-bulk.Rulings (supervisor step 3, 2026-10-04)
RAPIDPIPE_BATCH_RECLAIM_QUEUE; recorded in README "Running on Batch" and the module docstring.scheduler_metadata.batch.status_reasonholds only the job-level reason while the classifier also reads the last attempt's.rapid_docssystem/runs.md, "Attempts and retries") says a transient returns the unit to ready for a new attempt and does not name a queue, so this change does not contradict it. A sentence about the reclaim queue belongs there; that repo is outside this change's scope and the sentence is carried forward.Tests
tests/unit/test_launch.py: reclaim flag written for job- and attempt-levelHost EC2and not forCannot*/DockerTimeoutError*or when an exit code wins; unset/empty/equal → job queue with no lookup; reclaimed → reclaim queue; never reclaimed → job queue;resolve_joblessfinds a job on the reclaim queue, searches one queue when equal, and does not call a name/ARN double hit ambiguous.tests/db/test_launch.py(CI): reclaim → reclaim queue → exit 75 → still the reclaim queue; another unit of the run and the same unit id in another run stay on the job queue; a re-run of a re-run (twoseeded_from_unitgenerations) finds the reclaim; aCannot*transient resubmits to the job queue.python -m pytest tests/unit -q→ 1735 passed;scripts/check-public-safety.shexit 0.🤖 Generated with Claude Code