Skip to content

Change GraniteSwitch inheritance to MoEShared - #133

Draft
antonpibm wants to merge 6 commits into
mainfrom
feature/dehybridize-moeshared
Draft

antonpibm wants to merge 6 commits into
mainfrom
feature/dehybridize-moeshared

Conversation

@antonpibm

Copy link
Copy Markdown
Collaborator

No description provided.

Remove the hybrid (mamba/SSM) dependency from GraniteSwitch. The switch
model is attention-only and never constructs a mamba layer, so it does not
need the GraniteMoeHybrid family — GraniteMoeShared provides every class it
actually uses (shared MLP, MoE, RMSNorm, RoPE, PreTrainedModel base) minus
the mamba machinery.

HF backend:
- GraniteSwitchConfig now extends GraniteMoeSharedConfig; modeling and
  core/lora imports repointed to the granitemoeshared twins.
- layer_types / position_embedding_type become switch-owned attributes
  (the shared parent does not declare them, but internal readers still
  depend on them).

shared_intermediate_size fix: the shared parent defaults it to 0, which is
also the "no shared MLP" sentinel for pure sparse-MoE bases. The config now
resolves it itself instead of inheriting a magic default — explicit values
(including 0) are honored verbatim; when unset, dense resolves to
intermediate_size and pure MoE keeps 0. This is a compose-time decision
frozen into config.json; it also closes a latent bug where a bare dense
config silently inherited the old 1024 default. Guarded by new unit tests.

composer: granite_moe_hybrid_arch/_sr_arch renamed to
granite_moe_shared_arch/_sr_arch. The "granitemoehybrid" registry key is
retained (mapped to the shared arch) because real Granite 4.x dense
checkpoints are still typed granitemoehybrid upstream; a "granitemoeshared"
key is added alongside.

vLLM backend:
- Removed the vestigial IsHybrid / HasInnerState marker mixins (no hybrid
  contract was implemented).
- The two borrowed upstream classes (GraniteMoeMoE, GraniteMoeSharedMLP)
  now load via a version-tolerant helper that prefers the non-hybrid
  granitemoe / granitemoeshared modules and falls back to granitemoehybrid,
  so a single codebase works across the pinned vLLM versions.

Local CPU tests pass (unit, config sis-trap, composer arch skinning, HF
granite4 equivalence, HF forward/lora/multi-switch). vLLM and GPU
generation tests to run on the cluster.

Signed-off-by: antonp <antonp@il.ibm.com>
eval/gen_smoke.py loads a composed Granite Switch checkpoint in vLLM and
generates the same question with the base path and with each adapter's
control token, printing the outputs so a reviewer can confirm the
de-hybridized (granitemoeshared) backend both loads and routes adapters.
Used by the Vela validation job (vela_yamls/dehybridize_vllm_gen.yaml).

Signed-off-by: antonp <antonp@il.ibm.com>
Signed-off-by: antonp <antonp@il.ibm.com>
@codecov-commenter

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@antonpibm

Copy link
Copy Markdown
Collaborator Author

/gpu-test-multi

@github-actions

github-actions Bot commented Sep 15, 2026

Copy link
Copy Markdown

❌ GPU tests failed — vllm19-multi

No pytest summary — the suite did not start. See the log below.

Commit: 692c0e33d86c96a401097dc8cee5665d18359939
Full run & artifact log

Last 40 log lines
GPU test report — vllm19-multi
-------------------------------------------------------------------------
Commit:    692c0e33d86c96a401097dc8cee5665d18359939
Suite:     tests/unit/ tests/hf/ tests/composer/ tests/vllm/ tests/integration/
Deps:      dev (vllm 0.19.x)
GPU:       NVIDIA A100-SXM4-80GB

The test suite never started — the run failed during setup.
Redacted tail of the run log:

   + urllib3==2.7.0
   + uvicorn==0.47.0
   + uvloop==0.22.1
   + vllm==0.24.0
   + watchfiles==1.2.0
   + websockets==16.0
   + xgrammar==0.2.1
   + yarl==1.23.0
   + z3-solver==4.15.4.0
  torch 2.11.0+cu130 vllm 0.24.0
  FATAL expected vllm 0.19.x for group dev but got 0.24.0
  [05:40:11] <job> Failed
  [05:40:13] --- post-mortem: why the job ended (status and events only) ---
  [05:40:13] <job> status:
    QuotaReserved=True Resuming: Suspend is false
    ResourcesDeployed=False Resetting: Resources deleted for resetting <job>
    PodsReady=False Resetting: 
    Unhealthy=True FailedComponent: Found 1 failed components
    DeletingResources=False DeletionComplete: 
  [05:40:14]   pod does not exist — the job was never scheduled onto a node.
  [05:40:14]   For a multi-GPU suite the usual cause is waiting for capacity; see the
  [05:40:14]   <job> conditions above.
  [05:40:14] --- end post-mortem ---
  [05:40:14] cleanup: deleting <job> (exit=1)
  <job> "<job>" deleted

Verdict: FAILED before the test suite ran

@github-actions

github-actions Bot commented Sep 15, 2026

Copy link
Copy Markdown

❌ GPU tests failed — vllm20-multi

No pytest summary — the suite did not start. See the log below.

Commit: 692c0e33d86c96a401097dc8cee5665d18359939
Full run & artifact log

Last 40 log lines
GPU test report — vllm20-multi
-------------------------------------------------------------------------
Commit:    692c0e33d86c96a401097dc8cee5665d18359939
Suite:     tests/unit/ tests/hf/ tests/composer/ tests/vllm/ tests/integration/
Deps:      dev-vllm20 (vllm 0.20.x)
GPU:       NVIDIA A100-SXM4-80GB

The test suite never started — the run failed during setup.
Redacted tail of the run log:

  [05:40:03] <job> Failed
  [05:40:04] --- post-mortem: why the job ended (status and events only) ---
  [05:40:04] <job> status:
    QuotaReserved=True Resuming: Suspend is false
    ResourcesDeployed=True Resuming: Suspend is false
    PodsReady=True SufficientPodsReady: 1 pods running; 0 pods succeeded
    Unhealthy=False Resuming: Suspend is false
  [05:40:05] pod <job> phase and conditions:
    phase=Failed
    Initialized=True : 
    Ready=False PodFailed: 
    ContainersReady=False PodFailed: 
    PodScheduled=True : 
  [05:40:06] pod <job> container states:
    pytorch: ready=false restarts=0 state={"terminated":{"containerID":"cri-o://782217849f40f60b11680adb317a8b0df0e72fa6372ec8c38142bd8a9e294d64","exitCode":1,"finishedAt":"2026-09-15T05:39:54Z","reason":"Error","startedAt":"2026-09-15T05:38:16Z"}}
  [05:40:06] kubelet events for <job>:
  TIME                   TYPE     REASON           MESSAGE
  <nil>                  Normal   Scheduled        Successfully assigned security/<job> to dmf-nnnqh-gpu-worker-3-lxd5q
  2026-09-15T05:38:16Z   Normal   AddedInterface   Add eth0 [10.143.31.128/23] from ovn-kubernetes
  2026-09-15T05:38:16Z   Normal   Pulled           Container image "vllm/vllm-openai:latest" already present on machine
  2026-09-15T05:38:16Z   Normal   Created          Created container pytorch
  2026-09-15T05:38:16Z   Normal   Started          Started container pytorch
  [05:40:07] --- end post-mortem ---
  [05:40:07] cleanup: deleting <job> (exit=1)
  <job> "<job>" deleted

Verdict: FAILED before the test suite ran

Signed-off-by: antonp <antonp@il.ibm.com>
@antonpibm

Copy link
Copy Markdown
Collaborator Author

/gpu-test-multi

Signed-off-by: antonp <antonp@il.ibm.com>
Signed-off-by: antonp <antonp@il.ibm.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants