Repository navigation
test: prepare bounded Aliyun v3 performance campaigns - #204
Conversation
The v3 workspace benchmark had no maintained cloud runner lifecycle and the old ACK/Terraform tooling used stale deployment assumptions. Add an isolated Alibaba ECS/ESSD runner, private OSS evidence roundtrip, and a controller that preserves the original 15/235/240-minute window, frozen A/B inputs and existing full-verification oracles. Failures still destroy resources and audit exact disk, bucket, network, key-pair and runner identities. Production Rust is unchanged; OSS archives safe evidence, while the frozen native backend remains Local. Offline validation: 39 Python tests, Terraform fmt/validate and 8 mock-provider plans; official runner archive members and CLI response compatibility checked without registering runners or creating resources. Windows cannot create the archive symlinks; Linux CI exercises physical symlink creation. Plans and evidence stay outside the repository. Cloud bootstrap and performance execution require the separately reviewed launch plan.
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 93b3541caa
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| for name in environment)) | ||
| private_command(['/usr/bin/systemctl', 'daemon-reload'], payload) | ||
| private_command(['/usr/bin/systemctl', 'start', RUNNER_SERVICE], payload) | ||
| private_command(['/usr/bin/systemctl', 'is-active', '--quiet', RUNNER_SERVICE, STOP_TIMER], payload) |
There was a problem hiding this comment.
Check runner and timer activity separately
If the runner service exits immediately after systemctl start, this combined check still succeeds because the stop timer remains active. The systemctl is-active documentation specifies that, with multiple units, it “returns an exit code 0 if at least one is active”; consequently the controller can dispatch onto a dead runner and wait until the campaign cutoff. Check each unit in a separate command so both must be active.
Useful? React with 👍 / 👎.
| self.tools.api(f"repos/{REPO}/actions/workflows/{WORKFLOW}/dispatches", data={ | ||
| "ref": self.cfg["harness_ref"], "inputs": self.workflow_inputs()}) |
There was a problem hiding this comment.
Dispatch through an immutable harness reference
When harness_ref is a branch, as explicitly allowed by the example configuration, it can move after the preceding SHA check and before this separate dispatch request. GitHub then launches the workflow at the new head; discover_run() filters for the old head_sha, so it neither associates nor promptly cancels that run even though the unreviewed harness can execute on the dedicated runner. Dispatch through an immutable ref whose target is enforced, rather than merely checking a mutable ref immediately beforehand.
Useful? React with 👍 / 👎.
The v3 workspace benchmark had no maintained cloud runner lifecycle, and the old ACK/Terraform setup used stale deployment assumptions. Add one disposable Alibaba ECS/ESSD Actions runner, private OSS evidence roundtrip, and a controller that preserves the original session window, frozen A/B inputs, full-verification oracles and process/mount ownership checks. Failure cleanup checks disks, buckets, networks, key pairs and runner registrations using exact campaign identities.
Production Rust is unchanged. The native backend remains Local; OSS archives safe benchmark evidence. No cloud resources, runner registrations or performance campaigns were started.
Validation: independent review of the exact source tree; Linux offline CI 37750397721 passed 39 Python tests, Terraform fmt/validate and 8 mock-provider plans. Linux tests create and resolve the pinned runner symlinks. Official runner archive members and real OSS CLI response shapes were checked locally. Commits have valid GitHub signatures. Plans, logs and evidence remain outside the repository; actual cloud bootstrap and performance execution await the launch plan.