Repository navigation
Refactor/multiswitch only and conversation ids - #131
Merged
yairallouche merged 17 commits intoSep 10, 2026
Merged
Conversation
The doc was a standalone styled HTML page, which meant it rendered as raw markup on GitHub and sat outside the pre-commit link validator: that hook only reads tracked .md, .ipynb and .py files, and .html is absent from its EXT_OK set, so every link pointing AT the page was unchecked too. As Markdown it renders in-place and its inbound links are now validated. Structure is preserved: all 10 sections, 18 tables, and every ASCII diagram and verbatim trace. Typographic punctuation in prose is ASCII (the box-drawing characters in the diagrams are kept, being load-bearing); the table of contents is now anchor-linked. Pointers repointed at the new extension: CLAUDE.md gotcha 10 src/granite_switch/vllm/switch/multi.py:361 the 188-bound comment tutorials/notebooks/hello_multiswitch.ipynb tutorials/notebooks/multi_turn_multiswitch.ipynb tutorials/notebooks/multiswitch_serving.ipynb Content corrections made while converting: - switch_type's default is "multi" as of the next commit on this branch, so the comparison table and the config-surface table say so. Added a short subsection explaining why the two create_switch() call sites keep a "single" getattr fallback: it only fires for a config.json predating c5e78c6, whose num_hidden_layers was inflated by 1 rather than 2. - Stale line references corrected against this branch: config.py rows in the config-surface table, modeling_granite_switch.py:214-216 -> :343-344, conversation.py:99 -> :110. Remaining line references are carried over from the HTML and were not re-audited; the doc header now says so rather than implying otherwise. Signed-off-by: noaa <noaa.kless@ibm.com>
…t_payload Both policies reuse previously-sent ids as a stable prefix, so both must travel as ids. requires_token_ids is now constant True and the chat_payload method is removed (it returned a messages body the chat endpoint re-renders server-side, unable to carry ids). The /v1/chat/completions endpoint itself is unaffected; callers that want it just do not route through Conversation. conversation.py also carries the branch's SingleSwitch-removal cleanup of the old switch_type='multi' guard. Signed-off-by: noaa <noaa.kless@ibm.com>
Adds _appendable (control-token position decides whether the new turn can be a delta) and _control_char_index. A turn whose control token lands inside the already-sent prefix -- a LoRA token at index 0, or an aLoRA invocation in an earlier message -- can no longer be a delta; instead of raising, build_prompt re-prefills (full render, history's control tokens dropped). _reprefill is generalized with a reason so budget and placement fallbacks log accurately. A genuinely non-append-only template (Granite 4.2 thinking truncation) still raises: the discriminator is whether a no-adapter render extends the prefix. _reject_lora_placement is removed; _explain_no_prefix reduced to the template case. No adapter-technology detection anywhere. Signed-off-by: noaa <noaa.kless@ibm.com>
RE_PREFILL now carries _re_base_ids/_re_base_text: the base-rendered history prefix already sent (control tokens dropped), reused verbatim next turn so the prefix cache hits. build_prompt renders history with adapter=None (answers are already control-token-free, so this is base by construction) and appends the current turn's delta; a turn's base form is only sent the turn after it, so reuse lags one turn. Turn 1, a LoRA turn, or a non-append-only template full-renders. Output is byte-identical to a from-scratch RE_PREFILL render -- reuse is a cache optimization, never a change to what the model sees. Signed-off-by: noaa <noaa.kless@ibm.com>
Confirms the bf16 counting-head budget guard is PRESERVE-only: RE_PREFILL demotes history to base, so 50 turns each carry at most the current turn's one control token and never approach MAX_RETAINED_CONTROL_TOKENS. (The PRESERVE unappendable->reprefill wiring landed with Task 2.) Signed-off-by: noaa <noaa.kless@ibm.com>
RE_PREFILL now reuses base-demoted ids (lag one turn); an unappendable turn (LoRA at position 0, or an aLoRA trigger in an earlier turn) re-prefills instead of raising; chat_payload is removed and requires_token_ids is always true. Rewrites the two-policies table, the delta logic, the API surface, the refusals->fallback section, and recomputes the PRESERVE matrix with exact CPU-verified id counts. The one remaining raise is a non-append-only template. Test-inventory counts refreshed. Signed-off-by: noaa <noaa.kless@ibm.com>
Section 8 rewritten: PRESERVE no longer refuses a LoRA or stale-invocation turn, it re-prefills (demos now assert reprefills increments rather than a raise); chat_payload and the switch_type='single' demo are gone; the one remaining constructor raise is config-omitted, and the one build_prompt raise is a non-append-only template. Removes the dropped --switch-type CLI flag and the switch_type persistence assertion from section 9 (SingleSwitch is gone), and updates the intro, RE_PREFILL id-reuse framing, and the 188-ceiling note (re-prefill, not raise). Reprefill claims are backed by the unit tests. Signed-off-by: noaa <noaa.kless@ibm.com>
SingleSwitch is deleted from src, but the doc still described it as a live engine. create_switch no longer dispatches (it returns MultiSwitch directly); switch_type is no longer a config parameter but a rejection marker in from_dict. Rewrites the create_switch snippet, the switch_type-default table (now the rejection behavior), the comparison table (SingleSwitch marked removed, deleted-file reference dropped), the config-surface table (switch_type row removed), the layer-offset comment, and the base-reset row (dropped the removed validate_base_reset_switch_type). Also fixes the section-6 delta pseudocode: LoRA/earlier-trigger re-prefills, only a non-append-only template raises. Test reference updated to test_single_switch_rejected.py. Signed-off-by: noaa <noaa.kless@ibm.com>
SingleSwitch's single base->adapter transition (±gain cumsum over one attention head) is superseded by MultiSwitch's Kerdock/DG coded memory, which routes arbitrarily many transitions per request, latest-wins, and supports return-to-base. SingleSwitch averaged competing control tokens and mis-routed, and had no mechanism to re-select base mid-sequence. Delete both backends' single.py and all SingleSwitch-only tests and helpers (test_single_switch*, single_switch_cases, sequences, test_sharpness_equivalence; test_token_exchange's vLLM copy). create_switch builds MultiSwitch directly (no engine dispatch); from_dict rejects SingleSwitch/legacy checkpoints (pinned by test_single_switch_rejected.py); MultiSwitch owns num_cache_layers == 2 (counting + memory), where SingleSwitch owned 1. Signed-off-by: noaa <noaa.kless@ibm.com>
Four tests were added or shaped on main after this branch forked, while SingleSwitch was still the default engine. Removing SingleSwitch changes what they exercise, so update them to the MultiSwitch-only world: - test_control_lut_refresh.py / test_token_exchange.py: drop the ``"single"`` arm from the LUT-refresh parametrization; create_switch only ever returns MultiSwitch now, so the single arm asserted a type that can no longer be built. - test_granitemoe_compose_e2e.py: the switch reserves SWITCH_CACHE_LAYERS (== 2, MultiSwitch's counting + memory slots) at the front, not 1. Assert ``NUM_LAYERS + SWITCH_CACHE_LAYERS``; the old ``+ 1`` was the SingleSwitch single-slot layout. - test_multi_audio_compose_e2e.py: drop the ``--switch-type`` CLI flag (removed with SingleSwitch) and the ``single`` compose leg, and stop asserting a ``switch_type`` config key the composer no longer writes. The behavior each test pins (LUT refresh, front-loaded cache slots, audio compose shipping a consistent control LUT) is unchanged. Signed-off-by: noaa <noaa.kless@ibm.com>
The Vela run surfaced ~34 GPU/E2E failures the CPU suite never exercised, all from SingleSwitch assumptions left in tests after the engine was removed: - MultiSwitch reserves 2 cache slots (counting + memory), not 1. Fix the stale `num_hidden_layers - 1` in _model_forward_tests.py (its sibling at line 384 was already -2) and rewrite the KV-cache setup to register BOTH switch attention modules (counting_attn/memory_attn) instead of a single `.attn`. Size test_sr_switch's config to 4 layers (2 switch + 2 decoder) and drop the inert switch_type="single". - `single_overrides` was removed from generation_models; _noneager_generation_tests now uses basic_overrides (== switch_overrides, MultiSwitch). - Drop `config.switch_type` reads (the attribute no longer exists): the redundant asserts in test_multi_switch_alora/mixed_tech (both already isinstance-check MultiSwitch), the switch_type field emitted by the conversation workers, and the asserts consuming it. Build-phase anti-vacuity checks that read the raw config now key off the durable `ms_code_m` marker instead. Verified: test_sr_switch passes on CPU (25); all edited files clean under ruff 0.9.0 lint + format. The vLLM/integration ones re-verify on the next Vela run. Signed-off-by: noaa <noaa.kless@ibm.com>
…ingleSwitch preview test_quantization loaded ibm-granite/granite-switch-4.1-3b-preview, which the branch's from_dict now correctly rejects: that published checkpoint (like the 8b and 30b previews) is a legacy pre-coded-engine build — no ms_code_m, no switch_type — so it cannot load as MultiSwitch. No MultiSwitch checkpoint is published anywhere on the Hub. Follow the same pattern as test_multi_switch_mixed_tech: gate on GRANITE_SWITCH_E2E_MODELS=1 and compose (warm-reuse) a real ~3B mixed checkpoint from rag + guardian under GRANITE_SWITCH_E2E_DIR/multi-mixed. That yields exactly the two adapters ADAPTER_TESTS exercises — answerability (aLoRA, from rag) and hallucination_detection (LoRA, from guardian) — and reuses the identical checkpoint mixed_tech already composes, so no extra compose cost on a full run. GPU/E2E-only; verifies on the next Vela run. Signed-off-by: noaa <noaa.kless@ibm.com>
…h preview test_pipeline_parallelism_equivalence hardcoded ibm-granite/granite-switch-4.1-3b-preview, the same legacy SingleSwitch checkpoint from_dict now rejects — so both PP=1 and PP=2 worker steps died at config load. PP only needs a loadable MultiSwitch checkpoint with an active adapter to compare PP=1 vs PP=2 token equivalence, so gate on GRANITE_SWITCH_E2E_MODELS=1 and compose (warm-reuse) the same multi-mixed checkpoint test_quantization and test_multi_switch_mixed_tech build. GPU/E2E-only (needs 2 GPUs); verifies on the next Vela run. Signed-off-by: noaa <noaa.kless@ibm.com>
…ed adapters A GPU probe showed the 8 test_adapter_activates failures were neither a quantization nor a switch bug: the checkpoint the fixture composed did not contain answerability or hallucination_detection at all (its adapter_names were guardian- family: factuality-correction/-detection, guardian-core, policy-guardrails). So apply_chat_template(adapter_name="answerability") matched nothing, the control token was never injected (base and adapter prompts identical), and the adapter was a no-op in both bf16 and 4-bit — base == adapter, the assertion fails. granitelib-rag carries BOTH tested adapters (answerability as aLoRA, hallucination_detection as LoRA), so compose from rag alone into a dedicated "quant-rag" dir. The dedicated dir matters: the shared "multi-mixed" path may already hold a checkpoint composed from other libraries whose names would not match ADAPTER_TESTS. GPU/E2E-only; verifies on the next Vela run. Signed-off-by: noaa <noaa.kless@ibm.com>
….to() Removing SingleSwitch makes MultiSwitch the only (and default) engine, so every composed checkpoint now carries the fp32 Kerdock codebook buffer. The composer casts the model with model.to(bfloat16) (compose_utils.py), which nn.Module._apply would take the codebook down with it — but __init__ rebuilds the buffer as fp32 on every load. A cast model then writes a bf16 codebook while a reload of it writes fp32, so the same model serializes to two different byte counts (a 262,144-byte gap: 2048x64 fp32 vs bf16). Masked before the SingleSwitch removal because compose defaulted to the codebook-free SingleSwitch, so test_save_load_compose.py's byte-idempotency check (TestPhase2_DoubleSerialization::test_file_content_matches) never exercised the codebook path. With MultiSwitch as the default it does, and fails. _apply upcasts the codebook back to fp32 after any cast, restoring idempotency; the forward path is unaffected (codebook entries +/-0.125 are exact in bf16 and feed an fp32 dest). Verified: test_file_content_matches passes on CPU; tests/hf/test_multi_switch.py 297 pass. Signed-off-by: noaa <noaa.kless@ibm.com>
…rity Five branch-touched files had formatting the pinned ruff (0.9.0) reformats, and the aligned multi_turn_multiswitch notebook carried cell outputs. Apply ruff-format and nbstripout so `pre-commit run --all-files` (the CI source of truth) is clean. Signed-off-by: noaa <noaa.kless@ibm.com>
Remove session/branch cruft that does not help a reader understand the code: - MULTISWITCH_EXPLAINED.md: drop the "Converted from HTML on branch ... read at 1c86ee9 ... not been re-audited" provenance note, the "on this branch" framing, and the commit/branch archaeology for the alora-invocation-tail rule (kept the fact: the code uses the character rule, no alora_invocation_tail() helper here). - Drop "N bugs fixed on this branch" narrative from the multiswitch test docstrings, keeping the timeless rationale (each guards an end-to-end boundary). - refresh_switch_control_lut: drop the removed-SingleSwitch / --switch-type history, keep why it is extracted (testable without a real compose). - Retarget dangling references to the deleted SingleSwitch test test_switch_e2e_compose.py to the MultiSwitch serving e2e tests / generic notes. Signed-off-by: noaa <noaa.kless@ibm.com>
noaakl
requested review from
antonpibm,
aviv1ron1,
freunda and
yairallouche
as code owners
September 9, 2026 13:04
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
Collaborator
|
/gpu-test-multi |
✅ GPU tests passed —
|
✅ GPU tests passed —
|
yairallouche
approved these changes
Sep 10, 2026
aviv1ron1
added a commit
that referenced
this pull request
Sep 10, 2026
…ture #131 removed SingleSwitch. test_granitemoe_audio_compose.py arrived from feature/moe-audio-support, which branched before that, so it still parametrized its fixture over ["single", "multi"] and asserted the switch class was SingleSwitch on the single arm -- 9 errors after the merge, every one a [single] case, while all 23 [multi] cases passed. Now SWITCH_TYPES = ["multi"] and the assertion is unconditional, which is what every other switch-parametrized file already does post-#131: test_control_lut_refresh.py:34, tests/hf/test_multi_switch.py:49, tests/vllm/test_multi_switch.py:52, tests/shared/gap_equivalence.py:15. So this converges on the established shape rather than inventing one. Nothing about the audio path changed -- the coverage that was single/multi parametrized is engine-independent (the marker's embedding and output rows, the control-LUT refresh, save/load survival), so dropping the removed engine loses no assertion. 23 passed. The wider audio tier is 128 passed / 53 skipped with one error that is this machine having no GPU ("Torch not compiled with CUDA enabled"), not a merge regression. Signed-off-by: aviv ron <rona@il.ibm.com>
yairallouche
pushed a commit
that referenced
this pull request
Sep 30, 2026
#130) * Default ASR to Granite Speech 5.0 TurboCTC, and raise transformers to 5.16 Re-applies the TurboCTC work onto public main by hand, rather than by cherry-picking the six commits it was developed as (09c52c2..9f1e5e8 on the staging branch). Public main has since absorbed #95, #104, #116 and #121, so several of the files had moved underneath the original patches. What changes, in five parts: * DEFAULT_ASR_MODEL_ID becomes ibm-granite/granite-speech-5.0-470m-turboctc, a 470M English conformer CTC encoder, replacing distil-whisper/distil-small.en. Being CTC it has no generate(), so decode kwargs (language/task) are dropped rather than forwarded, and transcripts are lowercase and unpunctuated. * ASR defaults retuned for a CTC backend: asr_self_chunks False and asr_chunk_length_s 120.0, since the encoder has no internal chunking and the HF pipeline's own CTC chunking mis-trims every seam (it rescales stride by inputs_to_logits_ratio, which this checkpoint does not publish). asr_device defaults to cuda and asr_dtype resolves to bfloat16 there. * transformers pinned >=5.5.1,<5.17.0 in the core package and >=5.16.0 on the audio extra, which is where granite_speech5_ctc landed. Only the audio path needs 5.16, so the requirement sits on the extra. * ATTENTION_LAYER_TYPES accepts both "attention" and "full_attention". transformers 5.16 renamed the value and rewrites it inside PreTrainedConfig.__init__, so comparing against the bare string silently dropped the attention LoRA target groups and left adapters with MLP targets only. * IsHybrid is no longer declared. It requires mamba-state hooks this model neither implements nor needs -- the composer normalizes every layer to attention, so a composed checkpoint has zero mamba layers. vLLM's escape hatch for exactly this shape compares the literal "attention", so the 5.16 rename flipped is_hybrid true and engine init began failing for the mamba dtype hook. Two places where main had moved and the original patch could not be taken as-is. Both are in compose_granite_switch.py's argparse help, and both would have been regressions if cherry-picked: * main established that --enable-audio is the only flag that switches audio on ("Requires --enable-audio; ignored without it"), replacing the older "Implies --enable-audio". The re-applied text keeps main's rule and carries only the TurboCTC facts across. * docs/AUDIO.md was hand-merged rather than overwritten, so main's confirmation from the Granite authors that <|unused_N|> ids are reserved survives alongside the new TurboCTC sections. uv.lock regenerated rather than patched: transformers 5.8.1 -> 5.16.1. Known limitation, carried over and unresolved: vLLM <=0.25 keys a layer-type table on the old "attention" spelling in granitemoehybrid.py, so in a venv built from this branch, serving a *stock* Granite 4.0 hybrid checkpoint raises KeyError: 'full_attention'. Composed Granite Switch checkpoints are unaffected -- they declare GraniteSwitchForCausalLM, which builds its own layers and never consults that table. vLLM fixed it in 0.26.0. Verified: ruff check and format clean; tests/unit/test_asr.py, test_config.py and test_config_edge_cases.py pass 90/90 (2 skipped for lack of a local vLLM). Signed-off-by: aviv ron <rona@il.ibm.com> * fix(tests): stop _patched_pipeline leaking its mock onto transformers.pipeline Closes #121. _patched_pipeline patched transformers.pipelines.pipeline before transformers.pipeline. mock.patch.__enter__ records the current value so it can restore it, and transformers is a lazy module: resolving transformers.pipeline when it is not yet cached on the top-level module goes through transformers.pipelines. With the submodule patched first, the second patch read back the mock, recorded it as "the original", and faithfully restored it on exit. The mock then stayed installed for the rest of the process, and load() resolves `from transformers import pipeline` at call time, so every later real transcription picked it up. One of the mocks in this file raises the "does not recognize this architecture" ValueError, which _unsupported_architecture_error converts into ImportError: transformers 5.16.0 cannot load the ASR model 'ibm-granite/granite-speech-5.0-470m-turboctc': its architecture requires transformers>=5.16 -- naming the installed version as too old for itself, from a GPU test that had nothing to do with the unit test that leaked. That false trail is the reason this is worth more than a one-line diff of explanation. Swapping the two patches is the whole fix: transformers.pipeline is now read while transformers.pipelines is still real, so both record and restore the real function. Why it went unnoticed: the leak only occurs when transformers.pipeline is not already cached on the top-level module, which depends on what else ran first. tests/unit/test_asr.py alone restores correctly; the full tests/unit/ directory leaks. CI always runs the full suite, so CI always leaked. TestPatchedPipelineRestores guards both attributes, and forces the lazy-resolve precondition with transformers.__dict__.pop("pipeline", None) rather than relying on collection order. Verified to fail on the old ordering (assert <Mock> is not <Mock>) and pass on the new one, so it is a real guard rather than a tautology. Verified on transformers 5.16.1 and 5.8.1: full tests/unit/ is 287 passed / 1448 skipped on both, and a probe asserting transformers.pipeline is the real function after the session passes on both. ruff check and format clean. Note this fixes the misleading diagnosis, not the GPU-tier failures it was masking: the leaked vLLM engine that starves later GPU tests is a separate teardown defect (#123) and is deliberately left for its own PR. Signed-off-by: aviv ron <rona@il.ibm.com> * fix(vllm): alias full_attention so vLLM <=0.25 can load a 5.16-written config Closes #122. transformers 5.16 renamed the layer type "attention" to "full_attention" and rewrites the value inside PreTrainedConfig.__init__, so it is now the only spelling a written config can carry -- writing the old one back does not help. vLLM <=0.25 keys its Granite-hybrid layer table on the old name: # vllm/model_executor/models/granitemoehybrid.py:319 ALL_DECODER_LAYER_TYPES = {"attention": ..., "mamba": ...} layer_class = ALL_DECODER_LAYER_TYPES[config.layer_types[layer_idx]] KeyError: 'full_attention' register() now aliases the new name onto the same layer class. Reproduced on both 0.19.1 and 0.20.2; upstream fixed it in 0.26.0 by adding the key themselves, and setdefault makes this a no-op there, so the block can be deleted on a version bump without touching anything else. Two things this repairs, which is why it is not only a test fix: * The vLLM equivalence tier -- 14 failures plus 19 nested. Those tests build their upstream reference with transformers, save_pretrained it, and hand the directory to vllm.LLM; the saved config now says full_attention and its architectures field routes to vLLM's stale class, so the *control group* never starts. GraniteSwitchForCausalLM was never affected: it builds layers through a closure and does not consult that table. * Serving a stock Granite 4.0 hybrid checkpoint. In a venv built from this branch -- transformers>=5.16 comes from our own audio extra, vLLM is pinned <=0.25 -- `vllm serve ibm-granite/granite-4.0-micro` crashed. Composed Granite Switch checkpoints did not; the published 4.2-30b carries layer_types: ["full_attention", ...] and serves correctly. Why it lives in register() and not a test fixture: the lookup runs inside vLLM's *spawned* engine-core process, so nothing patched in the parent survives. vLLM loads its plugins in that process during init (v1/engine/core.py calls load_general_plugins()), reading the vllm.general_plugins entry point we already declare. It is the only hook we own that executes there. Aliasing adds a name, not behaviour: both keys resolve to the identical class object, so the decoder layer built is the one that was built before the rename. The broad except is deliberate -- if a future vLLM moves the module or the dict, skipping silently costs us this KeyError again, whereas raising would take down every engine start. A larger alternative was considered and rejected as too big for this change: retargeting the whole equivalence tier off granitemoehybrid onto the families we actually ship. That gap is real -- granite (4.1/4.2/30b) has no equivalence coverage in any tier and granitemoe has HF-only -- but it is separate work. tests/vllm/test_plugin_registration.py asserts both keys are present and resolve to the same object, and that register() stays re-entrant as its docstring promises. It runs in-process rather than through the usual subprocess wrapper because register() creates no engine and so opens no CUDA context; it skips cleanly where vLLM is absent. The assertions hold on 0.26+ too, so the test does not need editing when the alias becomes redundant. Verified: ruff check and format clean over 210 files; tests/unit/ is 286 passed / 1448 skipped on transformers 5.16.1, unchanged. The alias itself only demonstrates on GPU -- the proof is the 14 equivalence failures going green. Not addressed here: #123, the leaked vLLM engine that starves later GPU tests (8 failures + 8 errors), which is its own PR; and #127's tokenizer-fetch flake. Also untouched is the second stale comparison, is_hybrid in vllm/config/model.py, still broken in 0.28.0 -- we sidestep it by not declaring IsHybrid. Signed-off-by: aviv ron <rona@il.ibm.com> * docs+tests(audio): fold in the MoE-audio documentation and its coverage Absorbs feature/moe-audio-support so it does not need a PR of its own: that branch carried no functional code, only documentation plus the tests that back it. Taken from staging/feature/moe-audio-support @ 8bf0bb3. What came across: * docs/AUDIO.md -- the audio cascade over a pure sparse MoE base. Hand-merged, not overwritten: this file had already diverged here (63+/25- from the TurboCTC edits, which touch the same sections). Verified afterwards that both sides survive -- the TurboCTC default, the transformers>=5.16 requirement and the 120s chunker window on one side, the granitemoe material on the other. * tests/composer/test_audio_marker_output_row.py -- the marker/reserved-row copy is now parametrized over MLP topology as well as embedding tying, taking it from 2 cases to 4 (untied/tied x dense/sparse-MoE). The fixup only ever touches embedding rows, so both topologies must behave identically; the sparse arm is what would notice if a shared-MLP-shaped assumption crept into the model construction it runs against. * tests/composer/test_granitemoe_audio_compose.py -- new, 9 tests over 4 classes: audio does not resurrect the shared MLP, does not widen the adapter surface and leaves no zero-width parameter; the control LUT goes stale on the marker and refresh fixes it idempotently and passes the validator; the marker output row; and survival across save/load. Deliberately NOT taken: a 5-line comment block in src/granite_switch/composer/tokenizer_setup.py noting that the <|unused_N|> convention also holds on granitemoe bases. True, but prose, and keeping src/ out of this commit makes the "no functional change" claim checkable rather than asserted -- `git show --stat` shows no src/ path at all. Verified: the two test files give 32 passed here; ruff check and format clean over 211 files. One thing a reviewer should not misread. The composer tier on this base reports pre-existing failures that predate this commit and are unrelated to it: TestGraniteMoeSR fails 6 of 8 even in isolation, in 0.37s, with "ValueError: not enough values to unpack (expected 5, got 3)". Public main is 15 commits behind staging and is missing #119, which changed composer return signatures and updated test_granitemoe_compose_e2e.py to match; the two are out of step here. It has gone unnoticed because the public repo's automatic CI runs only tests/unit/, so tests/composer/ is unchecked on every PR. #119 is expected to reach public main shortly, which resolves it. The test file imported above is based on #116, before that refactor, so it matches the signatures this base has. Signed-off-by: aviv ron <rona@il.ibm.com> * test(audio): drop the SingleSwitch arm from the MoE-audio compose fixture #131 removed SingleSwitch. test_granitemoe_audio_compose.py arrived from feature/moe-audio-support, which branched before that, so it still parametrized its fixture over ["single", "multi"] and asserted the switch class was SingleSwitch on the single arm -- 9 errors after the merge, every one a [single] case, while all 23 [multi] cases passed. Now SWITCH_TYPES = ["multi"] and the assertion is unconditional, which is what every other switch-parametrized file already does post-#131: test_control_lut_refresh.py:34, tests/hf/test_multi_switch.py:49, tests/vllm/test_multi_switch.py:52, tests/shared/gap_equivalence.py:15. So this converges on the established shape rather than inventing one. Nothing about the audio path changed -- the coverage that was single/multi parametrized is engine-independent (the marker's embedding and output rows, the control-LUT refresh, save/load survival), so dropping the removed engine loses no assertion. 23 passed. The wider audio tier is 128 passed / 53 skipped with one error that is this machine having no GPU ("Torch not compiled with CUDA enabled"), not a merge regression. Signed-off-by: aviv ron <rona@il.ibm.com> * Revert the transformers 5.16 adaptation; leave it to the version-bump PR This branch should carry only the net TurboCTC change. The 5.16 work done here belongs with the version bump instead (issue #84, origin/issue-84-support-0.28), so that branch can be merged in cleanly rather than conflicting with a parallel half-done version of the same adaptation. Reverted to main exactly: pyproject.toml both transformers pins (core ceiling raise, and audio extra >=5.16) uv.lock back to main's resolution vllm/__init__.py the full_attention alias vllm/granite_switch_model.py IsHybrid declared again tests/unit/test_config_edge_cases.py literal layer_types assertions tests/shared/granite4_equivalence.py layer_types back to "attention" tests/vllm/test_plugin_registration.py deleted (it only tested the alias) Two files mixed both concerns and were edited rather than reverted: * config.py loses ATTENTION_LAYER_TYPES and goes back to comparing the bare string. Its TurboCTC content (the asr_* defaults and docstrings) stays. * docs/AUDIO.md claimed the audio extra *requires* transformers>=5.16, which is no longer true. Reworded rather than deleted, because the requirement is real even while unpinned: the default model needs >=5.16, that is not pinned yet, #84 tracks it, and until then the cascade runs only with an --asr-model the installed transformers supports. Kept deliberately: * The _unsupported_architecture_error guard in asr.py. It is TurboCTC's own error handling, not part of the bump, and it is what makes this branch coherent unpinned -- an install below 5.16 gets an actionable ImportError naming the fix instead of transformers' generic "unrecognized architecture". * The _patched_pipeline mock-leak fix. Pre-existing bug on main, but it is what stops the new TurboCTC GPU tests reporting a false "transformers 5.16 too old" error, which would make the next GPU run unreadable. * The MoE-audio documentation and tests. No 5.16 dependency; they pass on these pins. Done as a forward revert rather than a history rewrite: no force-push, PR #130 keeps its review history, and `git diff main..HEAD` is what determines the merge result either way. Two things the version-bump PR still needs, which this revert removes from here and which its current state does not cover: * ATTENTION_LAYER_TYPES accepting both spellings. Required on any transformers >=5.16 regardless of vLLM version -- comparing against the bare string silently drops the attention LoRA target groups, leaving adapters with MLP targets only. * Dropping IsHybrid. vLLM's is_hybrid escape hatch still compares the literal "attention" in 0.28.0, so moving to 0.28 does not fix it; a composed checkpoint still fails engine init asking for a mamba dtype hook. Only the full_attention alias is genuinely obsolete under 0.28 -- upstream added that key in 0.26.0 -- so that one is gone for good. Note the bump PR currently pins transformers <5.16, which still excludes the version the default ASR model needs. As it stands, merging this branch after it would leave TurboCTC unable to load. Verified: ruff check and format clean over 200 files; tests/unit/ is 235 passed / 1448 skipped on the restored pins. Signed-off-by: aviv ron <rona@il.ibm.com> --------- Signed-off-by: aviv ron <rona@il.ibm.com>
samadarsh
added a commit
to samadarsh/granite-switch
that referenced
this pull request
Oct 9, 2026
CLAUDE.md still described the removed SingleSwitch: the project tree listed hf/switch/single.py and vllm/switch/single.py, the import example and test commands pointed at SingleSwitch files that no longer exist, and the architecture section described a trainable router. It now reflects the coded-memory MultiSwitch (the only engine since generative-computing#131), links docs/MULTISWITCH_EXPLAINED.md, and lists conversation.py, token_exchange.py, hf/switch/codes/, vllm/audio/ and tests/audio/. The dev install command used `--extra dev`, but `dev` is a dependency group, so it is now `--group dev`; the nonexistent tests/regression/ entry is dropped. docs/AUDIO.md said transformers>=5.16 was not pinned yet. The package has required it since generative-computing#130, so the caveat now says a normal install already has it and keeps the ImportError fallback note for older installs. Signed-off-by: samadarsh <samadarsh14@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.