Skip to content

Refactor/multiswitch only and conversation ids - #131

Merged
yairallouche merged 17 commits into
mainfrom
refactor/multiswitch-only-and-conversation-ids
Sep 10, 2026
Merged

yairallouche merged 17 commits into
mainfrom
refactor/multiswitch-only-and-conversation-ids

Conversation

@noaakl

@noaakl noaakl commented Sep 9, 2026

Copy link
Copy Markdown
Collaborator

No description provided.

The doc was a standalone styled HTML page, which meant it rendered as raw
markup on GitHub and sat outside the pre-commit link validator: that hook
only reads tracked .md, .ipynb and .py files, and .html is absent from its
EXT_OK set, so every link pointing AT the page was unchecked too. As
Markdown it renders in-place and its inbound links are now validated.

Structure is preserved: all 10 sections, 18 tables, and every ASCII
diagram and verbatim trace. Typographic punctuation in prose is ASCII
(the box-drawing characters in the diagrams are kept, being load-bearing);
the table of contents is now anchor-linked.

Pointers repointed at the new extension:
  CLAUDE.md                                    gotcha 10
  src/granite_switch/vllm/switch/multi.py:361  the 188-bound comment
  tutorials/notebooks/hello_multiswitch.ipynb
  tutorials/notebooks/multi_turn_multiswitch.ipynb
  tutorials/notebooks/multiswitch_serving.ipynb

Content corrections made while converting:

- switch_type's default is "multi" as of the next commit on this branch,
  so the comparison table and the config-surface table say so. Added a
  short subsection explaining why the two create_switch() call sites keep
  a "single" getattr fallback: it only fires for a config.json predating
  c5e78c6, whose num_hidden_layers was inflated by 1 rather than 2.

- Stale line references corrected against this branch: config.py rows in
  the config-surface table, modeling_granite_switch.py:214-216 ->
  :343-344, conversation.py:99 -> :110.

Remaining line references are carried over from the HTML and were not
re-audited; the doc header now says so rather than implying otherwise.

Signed-off-by: noaa <noaa.kless@ibm.com>
…t_payload

Both policies reuse previously-sent ids as a stable prefix, so both must
travel as ids. requires_token_ids is now constant True and the chat_payload
method is removed (it returned a messages body the chat endpoint re-renders
server-side, unable to carry ids). The /v1/chat/completions endpoint itself
is unaffected; callers that want it just do not route through Conversation.

conversation.py also carries the branch's SingleSwitch-removal cleanup of the
old switch_type='multi' guard.

Signed-off-by: noaa <noaa.kless@ibm.com>
Adds _appendable (control-token position decides whether the new turn can be
a delta) and _control_char_index. A turn whose control token lands inside the
already-sent prefix -- a LoRA token at index 0, or an aLoRA invocation in an
earlier message -- can no longer be a delta; instead of raising, build_prompt
re-prefills (full render, history's control tokens dropped). _reprefill is
generalized with a reason so budget and placement fallbacks log accurately.

A genuinely non-append-only template (Granite 4.2 thinking truncation) still
raises: the discriminator is whether a no-adapter render extends the prefix.
_reject_lora_placement is removed; _explain_no_prefix reduced to the template
case. No adapter-technology detection anywhere.

Signed-off-by: noaa <noaa.kless@ibm.com>
RE_PREFILL now carries _re_base_ids/_re_base_text: the base-rendered history
prefix already sent (control tokens dropped), reused verbatim next turn so the
prefix cache hits. build_prompt renders history with adapter=None (answers are
already control-token-free, so this is base by construction) and appends the
current turn's delta; a turn's base form is only sent the turn after it, so
reuse lags one turn. Turn 1, a LoRA turn, or a non-append-only template
full-renders. Output is byte-identical to a from-scratch RE_PREFILL render --
reuse is a cache optimization, never a change to what the model sees.

Signed-off-by: noaa <noaa.kless@ibm.com>
Confirms the bf16 counting-head budget guard is PRESERVE-only: RE_PREFILL
demotes history to base, so 50 turns each carry at most the current turn's one
control token and never approach MAX_RETAINED_CONTROL_TOKENS. (The PRESERVE
unappendable->reprefill wiring landed with Task 2.)

Signed-off-by: noaa <noaa.kless@ibm.com>
RE_PREFILL now reuses base-demoted ids (lag one turn); an unappendable turn
(LoRA at position 0, or an aLoRA trigger in an earlier turn) re-prefills
instead of raising; chat_payload is removed and requires_token_ids is always
true. Rewrites the two-policies table, the delta logic, the API surface, the
refusals->fallback section, and recomputes the PRESERVE matrix with exact
CPU-verified id counts. The one remaining raise is a non-append-only template.
Test-inventory counts refreshed.

Signed-off-by: noaa <noaa.kless@ibm.com>
Section 8 rewritten: PRESERVE no longer refuses a LoRA or stale-invocation
turn, it re-prefills (demos now assert reprefills increments rather than a
raise); chat_payload and the switch_type='single' demo are gone; the one
remaining constructor raise is config-omitted, and the one build_prompt raise
is a non-append-only template. Removes the dropped --switch-type CLI flag and
the switch_type persistence assertion from section 9 (SingleSwitch is gone),
and updates the intro, RE_PREFILL id-reuse framing, and the 188-ceiling note
(re-prefill, not raise). Reprefill claims are backed by the unit tests.

Signed-off-by: noaa <noaa.kless@ibm.com>
SingleSwitch is deleted from src, but the doc still described it as a live
engine. create_switch no longer dispatches (it returns MultiSwitch directly);
switch_type is no longer a config parameter but a rejection marker in
from_dict. Rewrites the create_switch snippet, the switch_type-default table
(now the rejection behavior), the comparison table (SingleSwitch marked
removed, deleted-file reference dropped), the config-surface table (switch_type
row removed), the layer-offset comment, and the base-reset row (dropped the
removed validate_base_reset_switch_type). Also fixes the section-6 delta
pseudocode: LoRA/earlier-trigger re-prefills, only a non-append-only template
raises. Test reference updated to test_single_switch_rejected.py.

Signed-off-by: noaa <noaa.kless@ibm.com>
SingleSwitch's single base->adapter transition (±gain cumsum over one
attention head) is superseded by MultiSwitch's Kerdock/DG coded memory,
which routes arbitrarily many transitions per request, latest-wins, and
supports return-to-base. SingleSwitch averaged competing control tokens
and mis-routed, and had no mechanism to re-select base mid-sequence.

Delete both backends' single.py and all SingleSwitch-only tests and
helpers (test_single_switch*, single_switch_cases, sequences,
test_sharpness_equivalence; test_token_exchange's vLLM copy). create_switch
builds MultiSwitch directly (no engine dispatch); from_dict rejects
SingleSwitch/legacy checkpoints (pinned by test_single_switch_rejected.py);
MultiSwitch owns num_cache_layers == 2 (counting + memory), where
SingleSwitch owned 1.

Signed-off-by: noaa <noaa.kless@ibm.com>
Four tests were added or shaped on main after this branch forked, while
SingleSwitch was still the default engine. Removing SingleSwitch changes what
they exercise, so update them to the MultiSwitch-only world:

- test_control_lut_refresh.py / test_token_exchange.py: drop the ``"single"``
  arm from the LUT-refresh parametrization; create_switch only ever returns
  MultiSwitch now, so the single arm asserted a type that can no longer be built.
- test_granitemoe_compose_e2e.py: the switch reserves SWITCH_CACHE_LAYERS (== 2,
  MultiSwitch's counting + memory slots) at the front, not 1. Assert
  ``NUM_LAYERS + SWITCH_CACHE_LAYERS``; the old ``+ 1`` was the SingleSwitch
  single-slot layout.
- test_multi_audio_compose_e2e.py: drop the ``--switch-type`` CLI flag (removed
  with SingleSwitch) and the ``single`` compose leg, and stop asserting a
  ``switch_type`` config key the composer no longer writes.

The behavior each test pins (LUT refresh, front-loaded cache slots, audio compose
shipping a consistent control LUT) is unchanged.

Signed-off-by: noaa <noaa.kless@ibm.com>
The Vela run surfaced ~34 GPU/E2E failures the CPU suite never exercised, all
from SingleSwitch assumptions left in tests after the engine was removed:

- MultiSwitch reserves 2 cache slots (counting + memory), not 1. Fix the stale
  `num_hidden_layers - 1` in _model_forward_tests.py (its sibling at line 384 was
  already -2) and rewrite the KV-cache setup to register BOTH switch attention
  modules (counting_attn/memory_attn) instead of a single `.attn`. Size
  test_sr_switch's config to 4 layers (2 switch + 2 decoder) and drop the inert
  switch_type="single".
- `single_overrides` was removed from generation_models; _noneager_generation_tests
  now uses basic_overrides (== switch_overrides, MultiSwitch).
- Drop `config.switch_type` reads (the attribute no longer exists): the redundant
  asserts in test_multi_switch_alora/mixed_tech (both already isinstance-check
  MultiSwitch), the switch_type field emitted by the conversation workers, and the
  asserts consuming it. Build-phase anti-vacuity checks that read the raw config
  now key off the durable `ms_code_m` marker instead.

Verified: test_sr_switch passes on CPU (25); all edited files clean under ruff 0.9.0
lint + format. The vLLM/integration ones re-verify on the next Vela run.

Signed-off-by: noaa <noaa.kless@ibm.com>
…ingleSwitch preview

test_quantization loaded ibm-granite/granite-switch-4.1-3b-preview, which the
branch's from_dict now correctly rejects: that published checkpoint (like the 8b
and 30b previews) is a legacy pre-coded-engine build — no ms_code_m, no
switch_type — so it cannot load as MultiSwitch. No MultiSwitch checkpoint is
published anywhere on the Hub.

Follow the same pattern as test_multi_switch_mixed_tech: gate on
GRANITE_SWITCH_E2E_MODELS=1 and compose (warm-reuse) a real ~3B mixed checkpoint
from rag + guardian under GRANITE_SWITCH_E2E_DIR/multi-mixed. That yields exactly
the two adapters ADAPTER_TESTS exercises — answerability (aLoRA, from rag) and
hallucination_detection (LoRA, from guardian) — and reuses the identical checkpoint
mixed_tech already composes, so no extra compose cost on a full run.

GPU/E2E-only; verifies on the next Vela run.

Signed-off-by: noaa <noaa.kless@ibm.com>
…h preview

test_pipeline_parallelism_equivalence hardcoded ibm-granite/granite-switch-4.1-3b-preview,
the same legacy SingleSwitch checkpoint from_dict now rejects — so both PP=1 and
PP=2 worker steps died at config load. PP only needs a loadable MultiSwitch
checkpoint with an active adapter to compare PP=1 vs PP=2 token equivalence, so
gate on GRANITE_SWITCH_E2E_MODELS=1 and compose (warm-reuse) the same multi-mixed
checkpoint test_quantization and test_multi_switch_mixed_tech build.

GPU/E2E-only (needs 2 GPUs); verifies on the next Vela run.

Signed-off-by: noaa <noaa.kless@ibm.com>
…ed adapters

A GPU probe showed the 8 test_adapter_activates failures were neither a
quantization nor a switch bug: the checkpoint the fixture composed did not contain
answerability or hallucination_detection at all (its adapter_names were guardian-
family: factuality-correction/-detection, guardian-core, policy-guardrails). So
apply_chat_template(adapter_name="answerability") matched nothing, the control
token was never injected (base and adapter prompts identical), and the adapter was
a no-op in both bf16 and 4-bit — base == adapter, the assertion fails.

granitelib-rag carries BOTH tested adapters (answerability as aLoRA,
hallucination_detection as LoRA), so compose from rag alone into a dedicated
"quant-rag" dir. The dedicated dir matters: the shared "multi-mixed" path may
already hold a checkpoint composed from other libraries whose names would not
match ADAPTER_TESTS.

GPU/E2E-only; verifies on the next Vela run.

Signed-off-by: noaa <noaa.kless@ibm.com>
….to()

Removing SingleSwitch makes MultiSwitch the only (and default) engine, so every
composed checkpoint now carries the fp32 Kerdock codebook buffer. The composer
casts the model with model.to(bfloat16) (compose_utils.py), which nn.Module._apply
would take the codebook down with it — but __init__ rebuilds the buffer as fp32 on
every load. A cast model then writes a bf16 codebook while a reload of it writes
fp32, so the same model serializes to two different byte counts (a 262,144-byte
gap: 2048x64 fp32 vs bf16).

Masked before the SingleSwitch removal because compose defaulted to the
codebook-free SingleSwitch, so test_save_load_compose.py's byte-idempotency check
(TestPhase2_DoubleSerialization::test_file_content_matches) never exercised the
codebook path. With MultiSwitch as the default it does, and fails. _apply upcasts
the codebook back to fp32 after any cast, restoring idempotency; the forward path
is unaffected (codebook entries +/-0.125 are exact in bf16 and feed an fp32 dest).

Verified: test_file_content_matches passes on CPU; tests/hf/test_multi_switch.py 297 pass.
Signed-off-by: noaa <noaa.kless@ibm.com>
…rity

Five branch-touched files had formatting the pinned ruff (0.9.0) reformats, and
the aligned multi_turn_multiswitch notebook carried cell outputs. Apply
ruff-format and nbstripout so `pre-commit run --all-files` (the CI source of
truth) is clean.

Signed-off-by: noaa <noaa.kless@ibm.com>
Remove session/branch cruft that does not help a reader understand the code:

- MULTISWITCH_EXPLAINED.md: drop the "Converted from HTML on branch ... read at
  1c86ee9 ... not been re-audited" provenance note, the "on this branch" framing,
  and the commit/branch archaeology for the alora-invocation-tail rule (kept the
  fact: the code uses the character rule, no alora_invocation_tail() helper here).
- Drop "N bugs fixed on this branch" narrative from the multiswitch test
  docstrings, keeping the timeless rationale (each guards an end-to-end boundary).
- refresh_switch_control_lut: drop the removed-SingleSwitch / --switch-type history,
  keep why it is extracted (testable without a real compose).
- Retarget dangling references to the deleted SingleSwitch test
  test_switch_e2e_compose.py to the MultiSwitch serving e2e tests / generic notes.

Signed-off-by: noaa <noaa.kless@ibm.com>
@codecov-commenter

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.55072% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
src/granite_switch/conversation.py 98.24% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@aviv1ron1

Copy link
Copy Markdown
Collaborator

/gpu-test-multi

@github-actions

github-actions Bot commented Sep 9, 2026

Copy link
Copy Markdown

✅ GPU tests passed — vllm19-multi

2852 passed, 66 skipped, 22 warnings in 11068.91s (3:04:28)

Commit: d434cc524b890121b3c4a92d7edcd11184802ac0
Full run & artifact log

Last 40 log lines

.venv/lib/python3.12/site-packages/torch/jit/_script.py:362: 14 warnings
  /tmp/granite-switch/.venv/lib/python3.12/site-packages/torch/jit/_script.py:362: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`.
    warnings.warn(

tests/composer/test_compose_e2e.py:83
  /tmp/granite-switch/tests/composer/test_compose_e2e.py:83: PytestUnknownMarkWarning: Unknown pytest.mark.xdist_group - is this a typo?  You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html
    @pytest.mark.xdist_group("compose_e2e")

tests/composer/test_multi_audio_compose_e2e.py:124
  /tmp/granite-switch/tests/composer/test_multi_audio_compose_e2e.py:124: PytestUnknownMarkWarning: Unknown pytest.mark.xdist_group - is this a typo?  You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html
    @pytest.mark.xdist_group("multi_audio_compose_e2e")

tests/composer/test_multi_audio_compose_e2e.py:130
  /tmp/granite-switch/tests/composer/test_multi_audio_compose_e2e.py:130: PytestUnknownMarkWarning: Unknown pytest.mark.xdist_group - is this a typo?  You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html
    @pytest.mark.xdist_group("multi_audio_compose_e2e")

tests/composer/test_upstream_files.py:218
  /tmp/granite-switch/tests/composer/test_upstream_files.py:218: PytestUnknownMarkWarning: Unknown pytest.mark.xdist_group - is this a typo?  You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html
    @pytest.mark.xdist_group("upstream_build_e2e")

tests/integration/test_hf_to_vllm_weights.py::TestMultiSwitchForwardEquivalence::test_forward_logit_equivalence
  /tmp/granite-switch/.venv/lib/python3.12/site-packages/flashinfer/gemm/kernels/grouped_gemm_masked_blackwell.py:2059: DeprecationWarning: tcgen05.OperandMajorMode is deprecated, use cute.nvgpu.OperandMajorMode instead
    a_major_mode: tcgen05.OperandMajorMode,

tests/integration/test_hf_to_vllm_weights.py::TestMultiSwitchForwardEquivalence::test_forward_logit_equivalence
  /tmp/granite-switch/.venv/lib/python3.12/site-packages/flashinfer/gemm/kernels/grouped_gemm_masked_blackwell.py:2061: DeprecationWarning: tcgen05.OperandMajorMode is deprecated, use cute.nvgpu.OperandMajorMode instead
    b_major_mode: tcgen05.OperandMajorMode,

-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
========= 2852 passed, 66 skipped, 22 warnings in 11068.91s (3:04:28) ==========
sys:1: DeprecationWarning: builtin type swigvarlink has no __module__ attribute
[rank0]:[W909 16:34:15.132002371 ProcessGroupNCCL.cpp:1553] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
===== ALL GPU TESTS PASSED =====
[16:34:25] <job> Succeeded
[16:34:27] verified: found success sentinel in pod log
[16:34:27] cleanup: deleting <job> (exit=0)
<job> "<job>" deleted

Verdict: PASSED

@github-actions

github-actions Bot commented Sep 9, 2026

Copy link
Copy Markdown

✅ GPU tests passed — vllm20-multi

2852 passed, 66 skipped, 24 warnings in 11971.23s (3:19:31)

Commit: d434cc524b890121b3c4a92d7edcd11184802ac0
Full run & artifact log

Last 40 log lines
tests/composer/test_compose_e2e.py:83
  /tmp/granite-switch/tests/composer/test_compose_e2e.py:83: PytestUnknownMarkWarning: Unknown pytest.mark.xdist_group - is this a typo?  You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html
    @pytest.mark.xdist_group("compose_e2e")

tests/composer/test_multi_audio_compose_e2e.py:124
  /tmp/granite-switch/tests/composer/test_multi_audio_compose_e2e.py:124: PytestUnknownMarkWarning: Unknown pytest.mark.xdist_group - is this a typo?  You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html
    @pytest.mark.xdist_group("multi_audio_compose_e2e")

tests/composer/test_multi_audio_compose_e2e.py:130
  /tmp/granite-switch/tests/composer/test_multi_audio_compose_e2e.py:130: PytestUnknownMarkWarning: Unknown pytest.mark.xdist_group - is this a typo?  You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html
    @pytest.mark.xdist_group("multi_audio_compose_e2e")

tests/composer/test_upstream_files.py:218
  /tmp/granite-switch/tests/composer/test_upstream_files.py:218: PytestUnknownMarkWarning: Unknown pytest.mark.xdist_group - is this a typo?  You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html
    @pytest.mark.xdist_group("upstream_build_e2e")

tests/integration/test_hf_to_vllm_weights.py::TestMultiSwitchForwardEquivalence::test_forward_logit_equivalence
  /tmp/granite-switch/.venv/lib/python3.12/site-packages/flashinfer/gemm/kernels/grouped_gemm_masked_blackwell.py:2059: DeprecationWarning: tcgen05.OperandMajorMode is deprecated, use cute.nvgpu.OperandMajorMode instead
    a_major_mode: tcgen05.OperandMajorMode,

tests/integration/test_hf_to_vllm_weights.py::TestMultiSwitchForwardEquivalence::test_forward_logit_equivalence
  /tmp/granite-switch/.venv/lib/python3.12/site-packages/flashinfer/gemm/kernels/grouped_gemm_masked_blackwell.py:2061: DeprecationWarning: tcgen05.OperandMajorMode is deprecated, use cute.nvgpu.OperandMajorMode instead
    b_major_mode: tcgen05.OperandMajorMode,

tests/integration/test_hf_to_vllm_weights.py::TestMultiSwitchForwardEquivalence::test_forward_logit_equivalence
tests/integration/test_hf_to_vllm_weights.py::TestMultiSwitchForwardEquivalence::test_forward_logit_equivalence
  /tmp/granite-switch/.venv/lib/python3.12/site-packages/flashinfer/gdn_kernels/blackwell/gated_delta_net_chunked.py:99: DeprecationWarning: tcgen05.OperandMajorMode is deprecated, use cute.nvgpu.OperandMajorMode instead
    from cutlass.cute.nvgpu.tcgen05 import OperandMajorMode

-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
========= 2852 passed, 66 skipped, 24 warnings in 11971.23s (3:19:31) ==========
sys:1: DeprecationWarning: builtin type swigvarlink has no __module__ attribute
[rank0]:[W909 16:48:46.041874951 ProcessGroupNCCL.cpp:1575] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
===== ALL GPU TESTS PASSED =====
[16:48:54] <job> Succeeded
[16:48:56] verified: found success sentinel in pod log
[16:48:56] cleanup: deleting <job> (exit=0)
<job> "<job>" deleted

Verdict: PASSED

@yairallouche
yairallouche merged commit 6013c7f into main Sep 10, 2026
6 checks passed
aviv1ron1 added a commit that referenced this pull request Sep 10, 2026
…ture

#131 removed SingleSwitch. test_granitemoe_audio_compose.py arrived from
feature/moe-audio-support, which branched before that, so it still parametrized
its fixture over ["single", "multi"] and asserted the switch class was
SingleSwitch on the single arm -- 9 errors after the merge, every one a [single]
case, while all 23 [multi] cases passed.

Now SWITCH_TYPES = ["multi"] and the assertion is unconditional, which is what
every other switch-parametrized file already does post-#131:
test_control_lut_refresh.py:34, tests/hf/test_multi_switch.py:49,
tests/vllm/test_multi_switch.py:52, tests/shared/gap_equivalence.py:15. So this
converges on the established shape rather than inventing one.

Nothing about the audio path changed -- the coverage that was single/multi
parametrized is engine-independent (the marker's embedding and output rows, the
control-LUT refresh, save/load survival), so dropping the removed engine loses no
assertion.

23 passed. The wider audio tier is 128 passed / 53 skipped with one error that is
this machine having no GPU ("Torch not compiled with CUDA enabled"), not a merge
regression.

Signed-off-by: aviv ron <rona@il.ibm.com>
yairallouche pushed a commit that referenced this pull request Sep 30, 2026
#130)

* Default ASR to Granite Speech 5.0 TurboCTC, and raise transformers to 5.16

Re-applies the TurboCTC work onto public main by hand, rather than by
cherry-picking the six commits it was developed as (09c52c2..9f1e5e8 on the
staging branch). Public main has since absorbed #95, #104, #116 and #121, so
several of the files had moved underneath the original patches.

What changes, in five parts:

  * DEFAULT_ASR_MODEL_ID becomes ibm-granite/granite-speech-5.0-470m-turboctc,
    a 470M English conformer CTC encoder, replacing distil-whisper/distil-small.en.
    Being CTC it has no generate(), so decode kwargs (language/task) are dropped
    rather than forwarded, and transcripts are lowercase and unpunctuated.
  * ASR defaults retuned for a CTC backend: asr_self_chunks False and
    asr_chunk_length_s 120.0, since the encoder has no internal chunking and the
    HF pipeline's own CTC chunking mis-trims every seam (it rescales stride by
    inputs_to_logits_ratio, which this checkpoint does not publish). asr_device
    defaults to cuda and asr_dtype resolves to bfloat16 there.
  * transformers pinned >=5.5.1,<5.17.0 in the core package and >=5.16.0 on the
    audio extra, which is where granite_speech5_ctc landed. Only the audio path
    needs 5.16, so the requirement sits on the extra.
  * ATTENTION_LAYER_TYPES accepts both "attention" and "full_attention".
    transformers 5.16 renamed the value and rewrites it inside
    PreTrainedConfig.__init__, so comparing against the bare string silently
    dropped the attention LoRA target groups and left adapters with MLP targets
    only.
  * IsHybrid is no longer declared. It requires mamba-state hooks this model
    neither implements nor needs -- the composer normalizes every layer to
    attention, so a composed checkpoint has zero mamba layers. vLLM's escape
    hatch for exactly this shape compares the literal "attention", so the 5.16
    rename flipped is_hybrid true and engine init began failing for the mamba
    dtype hook.

Two places where main had moved and the original patch could not be taken as-is.
Both are in compose_granite_switch.py's argparse help, and both would have been
regressions if cherry-picked:

  * main established that --enable-audio is the only flag that switches audio on
    ("Requires --enable-audio; ignored without it"), replacing the older "Implies
    --enable-audio". The re-applied text keeps main's rule and carries only the
    TurboCTC facts across.
  * docs/AUDIO.md was hand-merged rather than overwritten, so main's confirmation
    from the Granite authors that <|unused_N|> ids are reserved survives
    alongside the new TurboCTC sections.

uv.lock regenerated rather than patched: transformers 5.8.1 -> 5.16.1.

Known limitation, carried over and unresolved: vLLM <=0.25 keys a layer-type
table on the old "attention" spelling in granitemoehybrid.py, so in a venv built
from this branch, serving a *stock* Granite 4.0 hybrid checkpoint raises
KeyError: 'full_attention'. Composed Granite Switch checkpoints are unaffected --
they declare GraniteSwitchForCausalLM, which builds its own layers and never
consults that table. vLLM fixed it in 0.26.0.

Verified: ruff check and format clean; tests/unit/test_asr.py, test_config.py and
test_config_edge_cases.py pass 90/90 (2 skipped for lack of a local vLLM).

Signed-off-by: aviv ron <rona@il.ibm.com>

* fix(tests): stop _patched_pipeline leaking its mock onto transformers.pipeline

Closes #121.

_patched_pipeline patched transformers.pipelines.pipeline before
transformers.pipeline. mock.patch.__enter__ records the current value so it can
restore it, and transformers is a lazy module: resolving transformers.pipeline
when it is not yet cached on the top-level module goes through
transformers.pipelines. With the submodule patched first, the second patch read
back the mock, recorded it as "the original", and faithfully restored it on exit.

The mock then stayed installed for the rest of the process, and load() resolves
`from transformers import pipeline` at call time, so every later real
transcription picked it up. One of the mocks in this file raises the "does not
recognize this architecture" ValueError, which _unsupported_architecture_error
converts into

    ImportError: transformers 5.16.0 cannot load the ASR model
      'ibm-granite/granite-speech-5.0-470m-turboctc': its architecture requires
      transformers>=5.16

-- naming the installed version as too old for itself, from a GPU test that had
nothing to do with the unit test that leaked. That false trail is the reason this
is worth more than a one-line diff of explanation.

Swapping the two patches is the whole fix: transformers.pipeline is now read
while transformers.pipelines is still real, so both record and restore the real
function.

Why it went unnoticed: the leak only occurs when transformers.pipeline is not
already cached on the top-level module, which depends on what else ran first.
tests/unit/test_asr.py alone restores correctly; the full tests/unit/ directory
leaks. CI always runs the full suite, so CI always leaked.

TestPatchedPipelineRestores guards both attributes, and forces the lazy-resolve
precondition with transformers.__dict__.pop("pipeline", None) rather than relying
on collection order. Verified to fail on the old ordering
(assert <Mock> is not <Mock>) and pass on the new one, so it is a real guard
rather than a tautology.

Verified on transformers 5.16.1 and 5.8.1: full tests/unit/ is 287 passed / 1448
skipped on both, and a probe asserting transformers.pipeline is the real function
after the session passes on both. ruff check and format clean.

Note this fixes the misleading diagnosis, not the GPU-tier failures it was
masking: the leaked vLLM engine that starves later GPU tests is a separate
teardown defect (#123) and is deliberately left for its own PR.

Signed-off-by: aviv ron <rona@il.ibm.com>

* fix(vllm): alias full_attention so vLLM <=0.25 can load a 5.16-written config

Closes #122.

transformers 5.16 renamed the layer type "attention" to "full_attention" and
rewrites the value inside PreTrainedConfig.__init__, so it is now the only
spelling a written config can carry -- writing the old one back does not help.
vLLM <=0.25 keys its Granite-hybrid layer table on the old name:

    # vllm/model_executor/models/granitemoehybrid.py:319
    ALL_DECODER_LAYER_TYPES = {"attention": ..., "mamba": ...}
    layer_class = ALL_DECODER_LAYER_TYPES[config.layer_types[layer_idx]]
    KeyError: 'full_attention'

register() now aliases the new name onto the same layer class. Reproduced on both
0.19.1 and 0.20.2; upstream fixed it in 0.26.0 by adding the key themselves, and
setdefault makes this a no-op there, so the block can be deleted on a version
bump without touching anything else.

Two things this repairs, which is why it is not only a test fix:

  * The vLLM equivalence tier -- 14 failures plus 19 nested. Those tests build
    their upstream reference with transformers, save_pretrained it, and hand the
    directory to vllm.LLM; the saved config now says full_attention and its
    architectures field routes to vLLM's stale class, so the *control group*
    never starts. GraniteSwitchForCausalLM was never affected: it builds layers
    through a closure and does not consult that table.
  * Serving a stock Granite 4.0 hybrid checkpoint. In a venv built from this
    branch -- transformers>=5.16 comes from our own audio extra, vLLM is pinned
    <=0.25 -- `vllm serve ibm-granite/granite-4.0-micro` crashed. Composed
    Granite Switch checkpoints did not; the published 4.2-30b carries
    layer_types: ["full_attention", ...] and serves correctly.

Why it lives in register() and not a test fixture: the lookup runs inside vLLM's
*spawned* engine-core process, so nothing patched in the parent survives. vLLM
loads its plugins in that process during init (v1/engine/core.py calls
load_general_plugins()), reading the vllm.general_plugins entry point we already
declare. It is the only hook we own that executes there.

Aliasing adds a name, not behaviour: both keys resolve to the identical class
object, so the decoder layer built is the one that was built before the rename.
The broad except is deliberate -- if a future vLLM moves the module or the dict,
skipping silently costs us this KeyError again, whereas raising would take down
every engine start.

A larger alternative was considered and rejected as too big for this change:
retargeting the whole equivalence tier off granitemoehybrid onto the families we
actually ship. That gap is real -- granite (4.1/4.2/30b) has no equivalence
coverage in any tier and granitemoe has HF-only -- but it is separate work.

tests/vllm/test_plugin_registration.py asserts both keys are present and resolve
to the same object, and that register() stays re-entrant as its docstring
promises. It runs in-process rather than through the usual subprocess wrapper
because register() creates no engine and so opens no CUDA context; it skips
cleanly where vLLM is absent. The assertions hold on 0.26+ too, so the test does
not need editing when the alias becomes redundant.

Verified: ruff check and format clean over 210 files; tests/unit/ is 286 passed /
1448 skipped on transformers 5.16.1, unchanged. The alias itself only
demonstrates on GPU -- the proof is the 14 equivalence failures going green.

Not addressed here: #123, the leaked vLLM engine that starves later GPU tests
(8 failures + 8 errors), which is its own PR; and #127's tokenizer-fetch flake.
Also untouched is the second stale comparison, is_hybrid in vllm/config/model.py,
still broken in 0.28.0 -- we sidestep it by not declaring IsHybrid.

Signed-off-by: aviv ron <rona@il.ibm.com>

* docs+tests(audio): fold in the MoE-audio documentation and its coverage

Absorbs feature/moe-audio-support so it does not need a PR of its own: that
branch carried no functional code, only documentation plus the tests that back
it. Taken from staging/feature/moe-audio-support @ 8bf0bb3.

What came across:

  * docs/AUDIO.md -- the audio cascade over a pure sparse MoE base. Hand-merged,
    not overwritten: this file had already diverged here (63+/25- from the
    TurboCTC edits, which touch the same sections). Verified afterwards that both
    sides survive -- the TurboCTC default, the transformers>=5.16 requirement and
    the 120s chunker window on one side, the granitemoe material on the other.
  * tests/composer/test_audio_marker_output_row.py -- the marker/reserved-row copy
    is now parametrized over MLP topology as well as embedding tying, taking it
    from 2 cases to 4 (untied/tied x dense/sparse-MoE). The fixup only ever
    touches embedding rows, so both topologies must behave identically; the sparse
    arm is what would notice if a shared-MLP-shaped assumption crept into the
    model construction it runs against.
  * tests/composer/test_granitemoe_audio_compose.py -- new, 9 tests over 4
    classes: audio does not resurrect the shared MLP, does not widen the adapter
    surface and leaves no zero-width parameter; the control LUT goes stale on the
    marker and refresh fixes it idempotently and passes the validator; the marker
    output row; and survival across save/load.

Deliberately NOT taken: a 5-line comment block in
src/granite_switch/composer/tokenizer_setup.py noting that the <|unused_N|>
convention also holds on granitemoe bases. True, but prose, and keeping src/
out of this commit makes the "no functional change" claim checkable rather than
asserted -- `git show --stat` shows no src/ path at all.

Verified: the two test files give 32 passed here; ruff check and format clean
over 211 files.

One thing a reviewer should not misread. The composer tier on this base reports
pre-existing failures that predate this commit and are unrelated to it:
TestGraniteMoeSR fails 6 of 8 even in isolation, in 0.37s, with
"ValueError: not enough values to unpack (expected 5, got 3)". Public main is 15
commits behind staging and is missing #119, which changed composer return
signatures and updated test_granitemoe_compose_e2e.py to match; the two are out
of step here. It has gone unnoticed because the public repo's automatic CI runs
only tests/unit/, so tests/composer/ is unchecked on every PR. #119 is expected
to reach public main shortly, which resolves it. The test file imported above is
based on #116, before that refactor, so it matches the signatures this base has.

Signed-off-by: aviv ron <rona@il.ibm.com>

* test(audio): drop the SingleSwitch arm from the MoE-audio compose fixture

#131 removed SingleSwitch. test_granitemoe_audio_compose.py arrived from
feature/moe-audio-support, which branched before that, so it still parametrized
its fixture over ["single", "multi"] and asserted the switch class was
SingleSwitch on the single arm -- 9 errors after the merge, every one a [single]
case, while all 23 [multi] cases passed.

Now SWITCH_TYPES = ["multi"] and the assertion is unconditional, which is what
every other switch-parametrized file already does post-#131:
test_control_lut_refresh.py:34, tests/hf/test_multi_switch.py:49,
tests/vllm/test_multi_switch.py:52, tests/shared/gap_equivalence.py:15. So this
converges on the established shape rather than inventing one.

Nothing about the audio path changed -- the coverage that was single/multi
parametrized is engine-independent (the marker's embedding and output rows, the
control-LUT refresh, save/load survival), so dropping the removed engine loses no
assertion.

23 passed. The wider audio tier is 128 passed / 53 skipped with one error that is
this machine having no GPU ("Torch not compiled with CUDA enabled"), not a merge
regression.

Signed-off-by: aviv ron <rona@il.ibm.com>

* Revert the transformers 5.16 adaptation; leave it to the version-bump PR

This branch should carry only the net TurboCTC change. The 5.16 work done here
belongs with the version bump instead (issue #84,
origin/issue-84-support-0.28), so that branch can be merged in cleanly rather
than conflicting with a parallel half-done version of the same adaptation.

Reverted to main exactly:

  pyproject.toml                          both transformers pins (core ceiling
                                          raise, and audio extra >=5.16)
  uv.lock                                 back to main's resolution
  vllm/__init__.py                        the full_attention alias
  vllm/granite_switch_model.py            IsHybrid declared again
  tests/unit/test_config_edge_cases.py    literal layer_types assertions
  tests/shared/granite4_equivalence.py    layer_types back to "attention"
  tests/vllm/test_plugin_registration.py  deleted (it only tested the alias)

Two files mixed both concerns and were edited rather than reverted:

  * config.py loses ATTENTION_LAYER_TYPES and goes back to comparing the bare
    string. Its TurboCTC content (the asr_* defaults and docstrings) stays.
  * docs/AUDIO.md claimed the audio extra *requires* transformers>=5.16, which is
    no longer true. Reworded rather than deleted, because the requirement is real
    even while unpinned: the default model needs >=5.16, that is not pinned yet,
    #84 tracks it, and until then the cascade runs only with an --asr-model the
    installed transformers supports.

Kept deliberately:

  * The _unsupported_architecture_error guard in asr.py. It is TurboCTC's own
    error handling, not part of the bump, and it is what makes this branch
    coherent unpinned -- an install below 5.16 gets an actionable ImportError
    naming the fix instead of transformers' generic "unrecognized architecture".
  * The _patched_pipeline mock-leak fix. Pre-existing bug on main, but it is what
    stops the new TurboCTC GPU tests reporting a false "transformers 5.16 too
    old" error, which would make the next GPU run unreadable.
  * The MoE-audio documentation and tests. No 5.16 dependency; they pass on these
    pins.

Done as a forward revert rather than a history rewrite: no force-push, PR #130
keeps its review history, and `git diff main..HEAD` is what determines the merge
result either way.

Two things the version-bump PR still needs, which this revert removes from here
and which its current state does not cover:

  * ATTENTION_LAYER_TYPES accepting both spellings. Required on any transformers
    >=5.16 regardless of vLLM version -- comparing against the bare string
    silently drops the attention LoRA target groups, leaving adapters with MLP
    targets only.
  * Dropping IsHybrid. vLLM's is_hybrid escape hatch still compares the literal
    "attention" in 0.28.0, so moving to 0.28 does not fix it; a composed
    checkpoint still fails engine init asking for a mamba dtype hook.

Only the full_attention alias is genuinely obsolete under 0.28 -- upstream added
that key in 0.26.0 -- so that one is gone for good.

Note the bump PR currently pins transformers <5.16, which still excludes the
version the default ASR model needs. As it stands, merging this branch after it
would leave TurboCTC unable to load.

Verified: ruff check and format clean over 200 files; tests/unit/ is 235 passed /
1448 skipped on the restored pins.

Signed-off-by: aviv ron <rona@il.ibm.com>

---------

Signed-off-by: aviv ron <rona@il.ibm.com>
samadarsh added a commit to samadarsh/granite-switch that referenced this pull request Oct 9, 2026
CLAUDE.md still described the removed SingleSwitch: the project tree listed
hf/switch/single.py and vllm/switch/single.py, the import example and test
commands pointed at SingleSwitch files that no longer exist, and the
architecture section described a trainable router. It now reflects the
coded-memory MultiSwitch (the only engine since generative-computing#131), links
docs/MULTISWITCH_EXPLAINED.md, and lists conversation.py, token_exchange.py,
hf/switch/codes/, vllm/audio/ and tests/audio/. The dev install command used
`--extra dev`, but `dev` is a dependency group, so it is now `--group dev`;
the nonexistent tests/regression/ entry is dropped.

docs/AUDIO.md said transformers>=5.16 was not pinned yet. The package has
required it since generative-computing#130, so the caveat now says a normal install already has
it and keeps the ImportError fallback note for older installs.

Signed-off-by: samadarsh <samadarsh14@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants