Skip to content

Consolidate benchmark shell scripts into shared lib + merged sweep script - #2

Merged
daedalus merged 6 commits into
masterfrom
copilot/update-benchmarks
Jul 15, 2026
Merged

Consolidate benchmark shell scripts into shared lib + merged sweep script#2
daedalus merged 6 commits into
masterfrom
copilot/update-benchmarks

Conversation

Copilot AI commented Jul 15, 2026

Copy link
Copy Markdown
Contributor

tools/bench.sh, tools/bench_sweep.sh, and tools/bench_sweep2.sh duplicated the same helper functions (cleanup_shm, extract, extract_ci, run_combo, verify_shm, check_coverage, run_with_retry) across ~650 lines, and bench_sweep2.sh re-ran many combos already covered by bench_sweep.sh.

Shared library

  • Extracted all common helpers into tools/lib/bench_common.sh, sourced by both remaining scripts
  • extract_ci now takes an explicit delimiter parameter (| for CSV rows, " " for table display) instead of callers post-processing the output

Merged sweep script

  • Folded tools/bench_sweep2.sh into tools/bench_sweep.sh (8 phases, 71 combos)
  • Deduplicated combos that became identical across the two scripts after a prior --meta-elo flag removal (s3/s4, gt1-gt3, f1-f3, etc.)
  • Preserved every genuinely unique combination from both originals (f16-f18, t8, z1-z4 variance-check phase)
  • Removed tools/bench_sweep2.sh

Bug fix

  • cleanup_shm piped grep output directly into while read; under set -o pipefail + set -e, this aborted the entire script whenever there were no stale SHM segments to clean (the common case). Fixed by capturing grep output into a variable before the loop.
  • run_with_retry now distinguishes "no log produced" (target crashed before startup) from "coverage failed to attach" for clearer retry diagnostics

Docs

  • Updated README.md and AGENTS.md to reflect the new tools/bench.sh / tools/bench_sweep.sh / tools/lib/bench_common.sh layout

Summary by Sourcery

Consolidate benchmark shell scripts into a shared library, merge and expand the feature-sweep benchmark, and align scripts/docs with the updated Elo/meta-elo behavior.

Bug Fixes:

  • Fix SHM cleanup to avoid aborting under set -e/pipefail when no segments are present.
  • Improve run_with_retry to distinguish missing log output from coverage-attachment failures for clearer diagnostics.

Enhancements:

  • Extract shared benchmarking helpers (SHM cleanup, metric extraction, coverage verification, retry logic, sweep combo runner) into tools/lib/bench_common.sh and source them from tools/bench.sh and tools/bench_sweep.sh.
  • Refine bench.sh and bench_sweep.sh configurations and sweep phases, removing obsolete meta-elo combinations, adding new candidate combos/variance checks, and increasing the number of top results shown.
  • Update Elo meta-scheduler reporting to match the consolidated --elo flag and adjust documentation to describe the new arbitration behavior and benchmark tooling layout.

@daedalus
daedalus marked this pull request as ready for review July 15, 2026 23:35
Copilot AI review requested due to automatic review settings July 15, 2026 23:35
@daedalus
daedalus merged commit 116e15e into master Jul 15, 2026
@daedalus
daedalus deleted the copilot/update-benchmarks branch July 15, 2026 23:36
@sourcery-ai

sourcery-ai Bot commented Jul 15, 2026

Copy link
Copy Markdown

Reviewer's Guide

Refactors the benchmarking shell scripts around a shared helper library, merges the two sweep scripts into a single, deduplicated benchmark sweep, fixes SHM cleanup and retry/coverage handling, aligns all benchmarks and docs with the removal of the deprecated --meta-elo flag, and updates Elo reporting to the new internal API.

Sequence diagram for run_with_retry coverage and SHM handling

sequenceDiagram
    participant BenchScript as bench.sh_or_bench_sweep.sh
    participant BenchLib as bench_common_sh
    participant Fuzzer as fuzzer_tool
    participant SHM as OS_SHM

    BenchScript->>BenchLib: run_with_retry(log, fuzz args)
    loop attempts up to BENCH_MAX_RETRIES
        BenchLib->>Fuzzer: python -m fuzzer_tool fuzz ...
        Fuzzer-->>BenchLib: write log
        alt [log is empty]
            BenchLib-->>BenchScript: "Run produced no log output" message
        else [log has content]
            BenchLib->>BenchLib: check_coverage(log, label)
            BenchLib->>BenchLib: verify_shm(log, label)
            BenchLib->>SHM: shmat(shm_id)
            SHM-->>BenchLib: bitmap bytes
            alt [bitmap has non-zero bytes]
                BenchLib-->>BenchScript: success
                Note over BenchLib,BenchScript: break loop
            else [coverage-blind run]
                BenchLib-->>BenchScript: "Coverage did not attach" message
            end
        end
        BenchLib->>BenchLib: cleanup_shm()
        BenchLib->>SHM: ipcs/ipcrm on orphaned segments
    end
    BenchLib-->>BenchScript: failure after BENCH_MAX_RETRIES
Loading

File-Level Changes

Change Details Files
Extract shared benchmark helpers into tools/lib/bench_common.sh and have bench.sh and bench_sweep.sh source it.
  • Move SHM cleanup, SHM verification, coverage checks, metric extraction, retry logic, and sweep combo runner into a new bench_common.sh library under tools/lib/
  • Update tools/bench.sh to source bench_common.sh instead of defining its own cleanup_shm, verify_shm, check_coverage, run_with_retry, extract, and extract_ci helpers inline.
  • Update tools/bench_sweep.sh to source bench_common.sh instead of inlining cleanup_shm, extract, extract_ci, and run_combo.
tools/lib/bench_common.sh
tools/bench.sh
tools/bench_sweep.sh
Improve robustness of SHM cleanup and retry/coverage handling in benchmark helpers.
  • Change cleanup_shm to capture grep output into a variable and iterate via a here-string to avoid failures under set -o pipefail when no SHM segments match.
  • Extend run_with_retry to treat an empty log file as a crash-before-startup condition, logging a distinct message before retrying.
  • Introduce BENCH_MAX_RETRIES environment variable (default 3) and use it in run_with_retry instead of a hard-coded MAX_RETRIES constant.
  • Centralize coverage verification via check_coverage/verify_shm in the shared library and use it from bench.sh.
tools/lib/bench_common.sh
tools/bench.sh
Merge bench_sweep2.sh into a single expanded bench_sweep.sh, deduplicating and extending feature combinations.
  • Remove tools/bench_sweep2.sh and port its unique combinations into tools/bench_sweep.sh.
  • Adjust benchmark phases to drop obsolete meta-elo variants and keep logically distinct combos, renaming where needed (e.g., f8_elo_bandit_markov_rep_shapley).
  • Add new combinations from the second sweep script such as f16–f18, t8, and z1–z4 variance-check runs.
  • Increase the final summary display from top 20 to top 40 edges-ranked results.
tools/bench_sweep.sh
tools/bench_sweep2.sh
Align all benchmark scripts and docs with the removal of the deprecated --meta-elo flag and new Elo behavior.
  • Remove --meta-elo from bench.sh configurations and update the printed configuration descriptions accordingly.
  • Remove --meta-elo from all combos in bench_sweep.sh and keep only the equivalent --elo-based variants.
  • Update AGENTS.md to describe Elo-based arbitration without a separate --meta-elo flag and to document the new benchmark script layout.
  • Update README.md benchmark examples to match the new enhanced configuration and to mention bench_sweep.sh and bench_common.sh.
  • Update docs/compose/reports/fuzzer-optimization-journey.md to drop --meta-elo from the example full feature-stack command.
  • Record the benchmark consolidation and meta-elo cleanup in docs/TODO.md.
tools/bench.sh
tools/bench_sweep.sh
AGENTS.md
README.md
docs/compose/reports/fuzzer-optimization-journey.md
docs/TODO.md
Fix Elo reporting to use the current internal Elo flag.
  • Change the meta-scheduler strategy ranking guard from f._use_meta_elo to f._use_elo so reports work with the consolidated Elo implementation.
src/fuzzer_tool/services/report.py
Adjust bench.sh reporting and CI extraction to use the new shared helper signature.
  • Remove local extract_ci implementation in bench.sh and call the shared helper with an explicit delimiter argument.
  • Change crash CI extraction in bench.sh to pass a space delimiter so the values can be printed directly in the summary table.
tools/bench.sh

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've found 2 issues, and left some high level feedback:

  • In bench_common.sh, verify_shm uses python3 while run_with_retry and run_combo use python; consider standardizing on one interpreter (python3 -m fuzzer_tool) to avoid environment-dependent failures.
  • run_with_retry treats both "no log produced" and explicit coverage failures as the same retry path and always prints "Coverage did not attach"; you might want to branch the messaging so the final failure reason clearly distinguishes between startup crashes and SHM-attachment issues.
Prompt for AI Agents
Please address the comments from this code review:

## Overall Comments
- In `bench_common.sh`, `verify_shm` uses `python3` while `run_with_retry` and `run_combo` use `python`; consider standardizing on one interpreter (`python3 -m fuzzer_tool`) to avoid environment-dependent failures.
- `run_with_retry` treats both "no log produced" and explicit coverage failures as the same retry path and always prints "Coverage did not attach"; you might want to branch the messaging so the final failure reason clearly distinguishes between startup crashes and SHM-attachment issues.

## Individual Comments

### Comment 1
<location path="tools/lib/bench_common.sh" line_range="16" />
<code_context>
+    # Capture matching SHM IDs into a variable first: under `set -o pipefail`,
+    # piping straight into `while read` would abort the script (via `set -e`)
+    # whenever grep finds no matches (the common case with no stale segments).
+    shmids=$(ipcs -m 2>/dev/null | grep "$(whoami)" | awk '{print $2}' || true)
+    if [[ -n "$shmids" ]]; then
+        while read -r shmid; do
</code_context>
<issue_to_address>
**suggestion (bug_risk):** User matching in SHM cleanup can accidentally match other usernames that contain the current username as a substring.

`grep "$(whoami)"` matches any line where the username appears as a substring (e.g., `foo` also matches `foobar`), so it can select SHM segments owned by other users. To restrict matches to the owner field, use something like:

- `awk '$3 == "'"$(whoami)"'" {print $2}'`, or
- `grep -E "^[^ ]+ +[^ ]+ +$(whoami) " | awk '{print $2}'`

so that only segments owned by the current user are removed.

Suggested implementation:

```
    local before shmids
    before=$(ipcs -m 2>/dev/null | awk -v user="$(whoami)" '$3 == user' | wc -l)
    # Capture matching SHM IDs into a variable first: under `set -o pipefail`,
    # piping straight into `while read` would abort the script (via `set -e`)
    # whenever the filter finds no matches (the common case with no stale segments).
    shmids=$(ipcs -m 2>/dev/null | awk -v user="$(whoami)" '$3 == user {print $2}')

```

```
    local after
    after=$(ipcs -m 2>/dev/null | awk -v user="$(whoami)" '$3 == user' | wc -l)

```
</issue_to_address>

### Comment 2
<location path="tools/lib/bench_common.sh" line_range="123-129" />
<code_context>
+# Runs `python -m fuzzer_tool "$@"`, verifying coverage attached; retries
+# on coverage-blind runs up to MAX_RETRIES (default 3) with SHM cleanup
+# between attempts.
+BENCH_MAX_RETRIES="${BENCH_MAX_RETRIES:-3}"
+
+run_with_retry() {
</code_context>
<issue_to_address>
**suggestion:** Consider validating `BENCH_MAX_RETRIES` to ensure it is a positive integer before using it in numeric comparisons.

Because `BENCH_MAX_RETRIES` is now environment-configurable, non-numeric or zero/negative values (e.g. `BENCH_MAX_RETRIES=foo` or `0`) can make `[[ $attempt -le $BENCH_MAX_RETRIES ]]` fail or behave unpredictably under `set -euo pipefail`. Consider normalizing the value once (e.g. default to 3 when unset or not a positive integer) so the retry loop stays robust to misconfiguration.

```suggestion
# ── Run with retry ────────────────────────────────────────────────────
# Runs `python -m fuzzer_tool "$@"`, verifying coverage attached; retries
# on coverage-blind runs up to MAX_RETRIES (default 3) with SHM cleanup
# between attempts.
BENCH_MAX_RETRIES="${BENCH_MAX_RETRIES:-3}"
# Normalize BENCH_MAX_RETRIES to a positive integer; fall back to 3 on
# invalid or non-positive values to keep the retry loop robust.
if ! [[ "$BENCH_MAX_RETRIES" =~ ^[1-9][0-9]*$ ]]; then
    echo "[!] Invalid BENCH_MAX_RETRIES='$BENCH_MAX_RETRIES'; using default of 3" >&2
    BENCH_MAX_RETRIES=3
fi

run_with_retry() {
```
</issue_to_address>

Sourcery is free for open source - if you like our reviews please consider sharing them ✨
Help me be more useful! Please click 👍 or 👎 on each comment and I'll use the feedback to improve your reviews.

Comment thread tools/lib/bench_common.sh
# Capture matching SHM IDs into a variable first: under `set -o pipefail`,
# piping straight into `while read` would abort the script (via `set -e`)
# whenever grep finds no matches (the common case with no stale segments).
shmids=$(ipcs -m 2>/dev/null | grep "$(whoami)" | awk '{print $2}' || true)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

suggestion (bug_risk): User matching in SHM cleanup can accidentally match other usernames that contain the current username as a substring.

grep "$(whoami)" matches any line where the username appears as a substring (e.g., foo also matches foobar), so it can select SHM segments owned by other users. To restrict matches to the owner field, use something like:

  • awk '$3 == "'"$(whoami)"'" {print $2}', or
  • grep -E "^[^ ]+ +[^ ]+ +$(whoami) " | awk '{print $2}'

so that only segments owned by the current user are removed.

Suggested implementation:

    local before shmids
    before=$(ipcs -m 2>/dev/null | awk -v user="$(whoami)" '$3 == user' | wc -l)
    # Capture matching SHM IDs into a variable first: under `set -o pipefail`,
    # piping straight into `while read` would abort the script (via `set -e`)
    # whenever the filter finds no matches (the common case with no stale segments).
    shmids=$(ipcs -m 2>/dev/null | awk -v user="$(whoami)" '$3 == user {print $2}')

    local after
    after=$(ipcs -m 2>/dev/null | awk -v user="$(whoami)" '$3 == user' | wc -l)

Comment thread tools/lib/bench_common.sh
Comment on lines +123 to +129
# ── Run with retry ────────────────────────────────────────────────────
# Runs `python -m fuzzer_tool "$@"`, verifying coverage attached; retries
# on coverage-blind runs up to MAX_RETRIES (default 3) with SHM cleanup
# between attempts.
BENCH_MAX_RETRIES="${BENCH_MAX_RETRIES:-3}"

run_with_retry() {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

suggestion: Consider validating BENCH_MAX_RETRIES to ensure it is a positive integer before using it in numeric comparisons.

Because BENCH_MAX_RETRIES is now environment-configurable, non-numeric or zero/negative values (e.g. BENCH_MAX_RETRIES=foo or 0) can make [[ $attempt -le $BENCH_MAX_RETRIES ]] fail or behave unpredictably under set -euo pipefail. Consider normalizing the value once (e.g. default to 3 when unset or not a positive integer) so the retry loop stays robust to misconfiguration.

Suggested change
# ── Run with retry ────────────────────────────────────────────────────
# Runs `python -m fuzzer_tool "$@"`, verifying coverage attached; retries
# on coverage-blind runs up to MAX_RETRIES (default 3) with SHM cleanup
# between attempts.
BENCH_MAX_RETRIES="${BENCH_MAX_RETRIES:-3}"
run_with_retry() {
# ── Run with retry ────────────────────────────────────────────────────
# Runs `python -m fuzzer_tool "$@"`, verifying coverage attached; retries
# on coverage-blind runs up to MAX_RETRIES (default 3) with SHM cleanup
# between attempts.
BENCH_MAX_RETRIES="${BENCH_MAX_RETRIES:-3}"
# Normalize BENCH_MAX_RETRIES to a positive integer; fall back to 3 on
# invalid or non-positive values to keep the retry loop robust.
if ! [[ "$BENCH_MAX_RETRIES" =~ ^[1-9][0-9]*$ ]]; then
echo "[!] Invalid BENCH_MAX_RETRIES='$BENCH_MAX_RETRIES'; using default of 3" >&2
BENCH_MAX_RETRIES=3
fi
run_with_retry() {

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR consolidates the benchmarking tooling by extracting duplicated shell helpers into a shared library, merges the second sweep script into the primary sweep, and aligns docs/reporting with the consolidation of the former --meta-elo behavior into --elo.

Changes:

  • Added tools/lib/bench_common.sh and updated tools/bench.sh / tools/bench_sweep.sh to source shared SHM cleanup, coverage verification, and log-extraction helpers.
  • Merged tools/bench_sweep2.sh scenarios into tools/bench_sweep.sh and removed tools/bench_sweep2.sh.
  • Updated reporting/docs to stop referencing removed meta-elo fields/flags and document the new benchmark layout.

Reviewed changes

Copilot reviewed 9 out of 11 changed files in this pull request and generated 2 comments.

Show a summary per file
File Description
tools/lib/bench_common.sh New shared benchmark helper library (SHM cleanup, extraction, coverage verification, retry, sweep runner).
tools/bench.sh Switched to shared helper library; removed --meta-elo usage; updated CI extraction formatting.
tools/bench_sweep.sh Switched to shared helper library; merged/expanded sweep phases; removed --meta-elo combos; increased results output.
tools/bench_sweep2.sh Removed (folded into tools/bench_sweep.sh).
src/fuzzer_tool/services/report.py Fixed report gating to use _use_elo instead of removed _use_meta_elo.
README.md Updated benchmark configuration docs and referenced shared helper library.
docs/TODO.md Recorded benchmark consolidation and the report.py _use_meta_elo fix.
docs/compose/reports/fuzzer-optimization-journey.md Updated example command to remove --meta-elo.
AGENTS.md Updated tools tree and documented Elo behavior without separate --meta-elo.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread tools/lib/bench_common.sh
Comment on lines +73 to +90
local has_data
has_data=$(python3 -c "
import ctypes, ctypes.util
libc = ctypes.CDLL(ctypes.util.find_library('c') or 'libc.so.6', use_errno=True)
libc.shmat.restype = ctypes.c_void_p
ptr = libc.shmat($shm_id, None, 0)
if ptr is None or ptr == -1:
print('FAIL')
else:
size = 4096 # default map size
bitmap = (ctypes.c_uint8 * size).from_address(ptr)
non_zero = sum(1 for i in range(size) if bitmap[i] != 0)
libc.shmdt(ptr)
if non_zero > 0:
print(f'OK:{non_zero}')
else:
print('EMPTY')
" 2>/dev/null)
Comment thread tools/lib/bench_common.sh

while [[ $attempt -le $BENCH_MAX_RETRIES ]]; do
echo "[*] Attempt $attempt/$BENCH_MAX_RETRIES..."
python -m fuzzer_tool "$@" 2>&1 | tee "$log"
daedalus added a commit that referenced this pull request Jul 16, 2026
Consolidate benchmark shell scripts into shared lib + merged sweep script
daedalus added a commit that referenced this pull request Jul 17, 2026
ENTROPY_HISTORY_MAX=200, ENTROPY_HISTORY_TRIM=100,
ENTROPY_WINDOW=4, ENTROPY_FLAT_THRESHOLD=0.001

All four findings verified:
- #2: zlib.crc32 returns deterministic int for LSH bucket keys ✓
- #3: grammar repeat bounds already clamped with max(hi, lo) ✓
- #4: edge_tracker uses zlib.crc32, not builtin hash() ✓
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants