You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Targets the production panic: ltsm: WASM abort called crash. Recovered LTSM exceptions were leaking WASM stack and heap until malloc aborted. This PR turns ordinary LTSM exceptions into errors and restores the WASM stack pointer after them. It also runs channel encrypt/decrypt on a disposable copy of the module, so their leaks can't exhaust the main runtime.
pkg/ltsm/embind.go: EM_JS and C++ throws panic with a new wasmException type. recoverError (deferred in callIndirect and Destroy) turns them into errors and restores g0; aborts and unexpected panics are re-panicked. noteChannelState marks the crypto copy stale when the main runtime touches Curve25519Key, E2EEKey, E2EEChannel or E2EEKeychain.
pkg/ltsm/runtime.go: channelCrypto runs E2EEChannelEncryptV1/V2 and DecryptV1/V2 on a shadow Runtime. The shadow is re-copied from the main module (memory, g0, emval table, fd state) at the start of the next channel call when it is stale, after any error or panic, or every 128 calls.
pkg/e2ee/manager.go: decrypts pick V1 or V2 from the first chunk's length (16 bytes = V2, anything else = V1). If that fails with a non-abort error, they fall back to the other version.
Tests: channel crypto leaves the main module untouched even when a call aborts, and version routing and fallback are covered.
Issues
All clear! No issues remaining. 🎉
7 issues already resolved
LTSM reports ordinary crypto failures (MAC/authentication failure, bad AES params) through the EM_JS "throw error" import, which panics, and checkDead now marks the runtime dead on any panic. One bad inbound ciphertext makes every later WASM call (including the V2→V1 decrypt fallback and request signing) fail until a reset. Only mark the runtime dead for ErrAbort, and turn EM_JS/C++ throws into plain errors as main did. (fixed by commit 2eb7ded)
Runner.reset clears keyStore, channelStore, goChannels and storageKey on the process-wide Runner. Every login's e2ee.Manager keeps those runner IDs (myKeyID, keyByRawID, channelByPair, groupKeys, groupSessionChannels) and never reloads them, so after one reset all E2EE sends and decrypts fail with unknown key/unknown channel for every user until restart or re-login. Rebuild the managers' state after a reset, or re-create WASM objects under the same IDs. (fixed by commit 2eb7ded)
E2EEChannelEncryptV2/DecryptV2 call writeStdString (which runs the module's malloc export directly) before CallMethod's dead check and outside callIndirect's checkDead. With encryptV2PanicSafe/decryptV2PanicSafe gone, an abort or panic there crashes the bridge instead of returning an error, and malloc still runs on an already-dead runtime. (fixed by commit 2eb7ded)
If reset() fails once, runnerErr keeps that error. A later successful reset then falls through to return globalRunner, runnerErr, so every caller gets the old error for good. Clear runnerErr when a reset succeeds, or return globalRunner, nil on that path. (fixed by commit 2eb7ded)
GetRunner reads rt.imp.dead holding only runnerMu, while WASM calls write it under Runner.mu, which go test -race will report. The runtimeDead hook is labeled a test seam, but no test uses it, so GetRunner/reset have no test coverage. (fixed by commit 2eb7ded)
Recovering an EM_JS/C++ exception unwinds WASM frames without restoring the stack pointer (Module.g0) or running C++ destructors, so each one leaks about 160–320 B of stack and up to ~5 KB of heap in the fixed 16 MB heap. After about 1.6k V2-decrypt exceptions (the V2→V1 fallback path), the next allocation calls _abort, for example in HmacDigest during GetSignature, which is the crash in this PR's description, so the PR does not fix it. Restore g0 on recovery, and avoid the throwing path or recycle the runtime before the heap runs out. (fixed by commit 72a3e0e)
Runtime.channelCrypto restores memory, g0, emval and fd state only after call() returns normally. If ErrAbort is re-panicked inside a channel encrypt/decrypt, the half-finished call's state stays in the module, and the runner's *PanicSafe wrappers then recover and keep using it. Run the restore in a defer so aborts are rolled back too. (fixed by commit 5152743)
I don't know if it's safe to recover WASM state like that, I presume we should probably unwind the whole bridge and relogin?
I don't think it's safe to reset the runtime here. Runner.reset() drops the encryption keys and channels but the e2ee managers still hold their IDs, so an abort like that would make messages unable to send or decrypt. checkDead also treats ordinary decrypt errors as fatal, which would break the LINE v2 (android & desktop) -> v1 (ios) decryption fallback
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Qwens attempt at fixing the following panic:
I don't know if it's safe to recover WASM state like that, I presume we should probably unwind the whole bridge and relogin?