Skip to content

fix(pricing): restore the DeepSeek v4 and gpt-5.6-codex repricing invariants (#1134) - #1577

Merged
iamtoruk merged 3 commits into
getagentseal:mainfrom
ozymandiashh:fix/1134-pricing-assertions
Sep 29, 2026
Merged

iamtoruk merged 3 commits into
getagentseal:mainfrom
ozymandiashh:fix/1134-pricing-assertions

Conversation

@ozymandiashh

Copy link
Copy Markdown
Collaborator

Closes #1134.

What this does

The 0.9.21 snapshot refresh caught the 2026-08-24 repricings mid-transition and relaxed two tests/models.test.ts invariants ("cannot be asserted until upstream syncs"). This PR restores both, now that upstream has synced, and fixes a bundler ordering defect the refresh exposed.

DeepSeek v4 — prefixed-equals-bare restored

LiteLLM now carries the official peak rates for both spellings (pro $1.32/$3.96 per million, flash $0.30/$1.20, cache-read $0.044/$0.006 — verified against api-docs.deepseek.com/quick_start/pricing, off-peak is half of peak). The hand-pinned MANUAL_ENTRIES for deepseek-v4-flash/deepseek-v4-pro (added while LiteLLM PR #27056 was still open) are removed: they shadowed upstream with the pre-sync promotional rates. deepseek/deepseek-v4-pro now equals deepseek-v4-pro, and the "official pricing" tests pin the current peak rates.

gpt-5.6-codex / -max — codex-equals-base restored

The codex SKUs are still absent from LiteLLM as of the 2026-09-29 refresh, so their MANUAL_ENTRIES remain — but now they mirror the repriced gpt-5.6 base row verbatim, tier block included ($4/$20 per million, 1.25x cache-write, >272k tier at 2x), per the same-generation pattern every codex id LiteLLM carries follows. The toEqual(snapshot['gpt-5.6']) assertion holds again, and if the base row moves, the mirror fails the test until refreshed. The long-context tier test now expects the codex SKUs to carry the base tier (previously undefined, because the flat manual tuple had no tier slot).

Bundler fix: reseller rows no longer shadow official rows

Pass 2 of scripts/bundle-litellm.mjs lets a prefixed row claim the slot its vendor-stripped key would answer. LiteLLM sorts reseller namespaces before vendor rows (openrouter/deepseek/deepseek-v4-pro at index 2167, deepseek/deepseek-v4-pro at 3375), so the openrouter resale row claimed the deepseek/deepseek-v4-pro slot first and the official row could never take it back (fillsOnly vetoes re-pricing) — the winner was decided purely by upstream JSON key order, and the namespaced id priced at the resale rate, ~40% under official peak. A prefixed row may now claim only a stripped slot that is not itself an upstream entry name; multi-segment ids that exist only as reseller rows (e.g. accounts/fireworks/models/...) keep resolving exactly as before.

Refreshed snapshot

Regenerated with npm run bundle-litellm: besides the deepseek/codex corrections, 24 rows moved by ordinary upstream drift (claude-sonnet-5-5 family added across providers, deepseek-v4.1-flash repriced, zai/glm-4.6 cache-write slot filled, two dropped mistral embed rows carried forward). Token-priced calls re-price on read, so no re-parse is needed.

Testing

  • tests/models.test.ts + tests/pricing-fallback-data.test.ts: 212 passed
  • tests/cli-deepseek-v4-pricing.test.ts: updated to the official peak rates (same observed token counts, recomputed dollars)
  • Full vitest run green apart from two pre-existing environment flakes (missing react-dom/server in the app workspace, one timing-sensitive electron retire test that passes standalone)

…ariants (getagentseal#1134)

The 0.9.21 snapshot caught the 2026-08-24 repricings mid-transition and
relaxed two tests/models.test.ts invariants. Upstream has since synced:

- DeepSeek v4: LiteLLM now carries the official peak rates for both the
  bare and the deepseek/-namespaced rows (pro 1.32/3.96 per million,
  flash 0.30/1.20), so the stale hand-pinned MANUAL_ENTRIES (from before
  PR #27056 merged) are removed and prefixed-equals-bare is asserted
  again.
- gpt-5.6-codex / -max: still absent from LiteLLM, so their
  MANUAL_ENTRIES remain but now mirror the repriced gpt-5.6 base row
  verbatim (4/20 per million, tier block included). codex-equals-base
  is asserted again, and the long-context tier test now expects the
  codex SKUs to carry the base tier.

The refresh also exposed a pass-2 ordering defect in
scripts/bundle-litellm.mjs: a reseller row like
openrouter/deepseek/deepseek-v4-pro could claim the empty
deepseek/deepseek-v4-pro slot before the official row was reached (the
winner was decided purely by upstream JSON key order), pricing the
namespaced id at the resale rate. A prefixed row may now claim only a
stripped slot that is not itself an upstream entry name; multi-segment
ids that exist only as reseller rows keep resolving as before.

Bundled snapshot regenerated: 24 rows moved by upstream drift
(claude-sonnet-5-5 family added, deepseek-v4.1-flash repriced,
zai/glm-4.6 cache-write filled), plus the deepseek/codex corrections.
Token-priced calls re-price on read; no re-parse is needed.
… shadow bug

- session-cache.ts: bump the cline-cli parse version so cached estimated
  costs (e.g. offline DeepSeek v4.1 flash) re-price on read, since
  cline-cli is in REPORTED_COST_PROVIDERS.
- CHANGELOG.md: note the bundler fix also reprices zai/glm-4.6 from the
  vercel_ai_gateway resale row to the official row, and moves the two
  mistral embed models to the fallback table.
- bundle-litellm.test.ts: add a fixture covering a reseller-prefixed row
  that sorts before the official row it would otherwise shadow.
@iamtoruk
iamtoruk merged commit 986d72c into getagentseal:main Sep 29, 2026
20 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

pricing: re-tighten gpt-5.6-codex and DeepSeek v4 snapshot assertions once LiteLLM syncs the 2026-08-24 repricings

2 participants