Skip to content

[Fix] Avoid CUDA sync in DeepSeek-V4 prefill token metadata - #5016

Open
adenzhou1350 wants to merge 1 commit into
InternLM:mainfrom
adenzhou1350:codex/v4-prefill-host-token-count
Open

adenzhou1350 wants to merge 1 commit into
InternLM:mainfrom
adenzhou1350:codex/v4-prefill-host-token-count

Avoid a CUDA scalar read when building V4 prefill metadata

b61129b
Select commit
Loading
Failed to load commit list.
Sign in for the full log view