Skip to content

launch: a Spot-reclaimed unit's further attempts go to RAPIDPIPE_BATCH_RECLAIM_QUEUE - #225

Merged
rusholme merged 1 commit into
rebuildfrom
reclaim-queue
Oct 5, 2026
Merged

rusholme merged 1 commit into
rebuildfrom
reclaim-queue

Conversation

@rusholme

@rusholme rusholme commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator

What

The launcher already records a Spot reclaim (FAILED, no container exit code, statusReason starting Host EC2) as transient and resubmits the unit as a fresh attempt, but always on the process-wide RAPIDPIPE_BATCH_JOB_QUEUE, so a reclaimed unit could go straight back to Spot. This adds an optional RAPIDPIPE_BATCH_RECLAIM_QUEUE:

  • reconcile marks an attempt lost to a host reclaim with scheduler_metadata.batch.reclaim = true (beside the job_queue it already records). Other transients (exit 75, Cannot*, DockerTimeoutError*) are not marked.
  • submit_unit (via queue_for_unit) sends every further attempt of a unit to the reclaim queue once any attempt of it, in the run or any run it was seeded from (units.seeded_from_unit, followed through every seed), was a reclaim. Sticky: a later exit 75 does not send it back to Spot. A never-reclaimed unit keeps the job queue.
  • Unset, empty, or equal to RAPIDPIPE_BATCH_JOB_QUEUE: the same submit_job call as before and no database lookup.
  • run reconcile --resolve-jobless searches both queues when they differ, counting distinct job ids (a queue's name and ARN are two spellings of one queue).

Env contract: RAPIDPIPE_BATCH_RECLAIM_QUEUE, when set and different from RAPIDPIPE_BATCH_JOB_QUEUE, receives every further attempt of a unit reclaimed in its run or a run it was seeded from; unset or equal, behaviour is unchanged.

Production stays unchanged: the processing-date loop (rapid_systems op-processing-date-loop.sh) will export both variables as rapid-queue-prompt (on-demand) in a follow-up PR there. A future Spot test sets only the job queue to rapid-queue-bulk.

Rulings (supervisor step 3, 2026-10-04)

  • Env name RAPIDPIPE_BATCH_RECLAIM_QUEUE; recorded in README "Running on Batch" and the module docstring.
  • The reclaim is marked at reconcile rather than re-matched at submit time, because scheduler_metadata.batch.status_reason holds only the job-level reason while the classifier also reads the last attempt's.
  • Sticky per unit, across seeded re-runs, because the intent is "shift to on-demand once reclaimed and never retry on Spot" (Ben, 2026-10-02).
  • Not chosen: an EventBridge+Lambda resubmit (bypasses the attempt record); compute-environment overflow Spot to on-demand (removed to pin lanes to quotas); a lane-to-queue map in the launcher (a bigger change than the gap).
  • The runs page (rapid_docs system/runs.md, "Attempts and retries") says a transient returns the unit to ready for a new attempt and does not name a queue, so this change does not contradict it. A sentence about the reclaim queue belongs there; that repo is outside this change's scope and the sentence is carried forward.

Tests

  • tests/unit/test_launch.py: reclaim flag written for job- and attempt-level Host EC2 and not for Cannot*/DockerTimeoutError* or when an exit code wins; unset/empty/equal → job queue with no lookup; reclaimed → reclaim queue; never reclaimed → job queue; resolve_jobless finds a job on the reclaim queue, searches one queue when equal, and does not call a name/ARN double hit ambiguous.
  • tests/db/test_launch.py (CI): reclaim → reclaim queue → exit 75 → still the reclaim queue; another unit of the run and the same unit id in another run stay on the job queue; a re-run of a re-run (two seeded_from_unit generations) finds the reclaim; a Cannot* transient resubmits to the job queue.
  • Local: python -m pytest tests/unit -q → 1735 passed; scripts/check-public-safety.sh exit 0.

🤖 Generated with Claude Code

…H_RECLAIM_QUEUE

reconcile marks an attempt lost to a host reclaim (statusReason "Host
EC2", job or last attempt) with scheduler_metadata.batch.reclaim.
submit_unit then sends every further attempt of that unit, in the run or
any run seeded from it, to RAPIDPIPE_BATCH_RECLAIM_QUEUE, so a reclaimed
unit moves to on-demand capacity and never retries on Spot. Any other
transient keeps the job queue. Unset, or equal to
RAPIDPIPE_BATCH_JOB_QUEUE, nothing changes and no lookup is made.
resolve_jobless searches both queues when they differ, counting distinct
job ids.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@rusholme
rusholme merged commit d6901dc into rebuild Oct 5, 2026
11 checks passed
@rusholme
rusholme deleted the reclaim-queue branch October 5, 2026 02:48
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant