fix(pricing): restore the DeepSeek v4 and gpt-5.6-codex repricing invariants (#1134) - #1577
Merged
iamtoruk merged 3 commits intoSep 29, 2026
Merged
Conversation
…ariants (getagentseal#1134) The 0.9.21 snapshot caught the 2026-08-24 repricings mid-transition and relaxed two tests/models.test.ts invariants. Upstream has since synced: - DeepSeek v4: LiteLLM now carries the official peak rates for both the bare and the deepseek/-namespaced rows (pro 1.32/3.96 per million, flash 0.30/1.20), so the stale hand-pinned MANUAL_ENTRIES (from before PR #27056 merged) are removed and prefixed-equals-bare is asserted again. - gpt-5.6-codex / -max: still absent from LiteLLM, so their MANUAL_ENTRIES remain but now mirror the repriced gpt-5.6 base row verbatim (4/20 per million, tier block included). codex-equals-base is asserted again, and the long-context tier test now expects the codex SKUs to carry the base tier. The refresh also exposed a pass-2 ordering defect in scripts/bundle-litellm.mjs: a reseller row like openrouter/deepseek/deepseek-v4-pro could claim the empty deepseek/deepseek-v4-pro slot before the official row was reached (the winner was decided purely by upstream JSON key order), pricing the namespaced id at the resale rate. A prefixed row may now claim only a stripped slot that is not itself an upstream entry name; multi-segment ids that exist only as reseller rows keep resolving as before. Bundled snapshot regenerated: 24 rows moved by upstream drift (claude-sonnet-5-5 family added, deepseek-v4.1-flash repriced, zai/glm-4.6 cache-write filled), plus the deepseek/codex corrections. Token-priced calls re-price on read; no re-parse is needed.
… shadow bug - session-cache.ts: bump the cline-cli parse version so cached estimated costs (e.g. offline DeepSeek v4.1 flash) re-price on read, since cline-cli is in REPORTED_COST_PROVIDERS. - CHANGELOG.md: note the bundler fix also reprices zai/glm-4.6 from the vercel_ai_gateway resale row to the official row, and moves the two mistral embed models to the fallback table. - bundle-litellm.test.ts: add a fixture covering a reseller-prefixed row that sorts before the official row it would otherwise shadow.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #1134.
What this does
The 0.9.21 snapshot refresh caught the 2026-08-24 repricings mid-transition and relaxed two
tests/models.test.tsinvariants ("cannot be asserted until upstream syncs"). This PR restores both, now that upstream has synced, and fixes a bundler ordering defect the refresh exposed.DeepSeek v4 — prefixed-equals-bare restored
LiteLLM now carries the official peak rates for both spellings (pro $1.32/$3.96 per million, flash $0.30/$1.20, cache-read $0.044/$0.006 — verified against api-docs.deepseek.com/quick_start/pricing, off-peak is half of peak). The hand-pinned
MANUAL_ENTRIESfordeepseek-v4-flash/deepseek-v4-pro(added while LiteLLM PR #27056 was still open) are removed: they shadowed upstream with the pre-sync promotional rates.deepseek/deepseek-v4-pronow equalsdeepseek-v4-pro, and the "official pricing" tests pin the current peak rates.gpt-5.6-codex / -max — codex-equals-base restored
The codex SKUs are still absent from LiteLLM as of the 2026-09-29 refresh, so their
MANUAL_ENTRIESremain — but now they mirror the repricedgpt-5.6base row verbatim, tier block included ($4/$20 per million, 1.25x cache-write, >272k tier at 2x), per the same-generation pattern every codex id LiteLLM carries follows. ThetoEqual(snapshot['gpt-5.6'])assertion holds again, and if the base row moves, the mirror fails the test until refreshed. The long-context tier test now expects the codex SKUs to carry the base tier (previously undefined, because the flat manual tuple had no tier slot).Bundler fix: reseller rows no longer shadow official rows
Pass 2 of
scripts/bundle-litellm.mjslets a prefixed row claim the slot its vendor-stripped key would answer. LiteLLM sorts reseller namespaces before vendor rows (openrouter/deepseek/deepseek-v4-proat index 2167,deepseek/deepseek-v4-proat 3375), so the openrouter resale row claimed thedeepseek/deepseek-v4-proslot first and the official row could never take it back (fillsOnly vetoes re-pricing) — the winner was decided purely by upstream JSON key order, and the namespaced id priced at the resale rate, ~40% under official peak. A prefixed row may now claim only a stripped slot that is not itself an upstream entry name; multi-segment ids that exist only as reseller rows (e.g.accounts/fireworks/models/...) keep resolving exactly as before.Refreshed snapshot
Regenerated with
npm run bundle-litellm: besides the deepseek/codex corrections, 24 rows moved by ordinary upstream drift (claude-sonnet-5-5 family added across providers,deepseek-v4.1-flashrepriced,zai/glm-4.6cache-write slot filled, two dropped mistral embed rows carried forward). Token-priced calls re-price on read, so no re-parse is needed.Testing
tests/models.test.ts+tests/pricing-fallback-data.test.ts: 212 passedtests/cli-deepseek-v4-pricing.test.ts: updated to the official peak rates (same observed token counts, recomputed dollars)vitest rungreen apart from two pre-existing environment flakes (missingreact-dom/serverin the app workspace, one timing-sensitive electron retire test that passes standalone)