fix(openclaw): read the per-agent sqlite transcript store the 2026-09-01 migration writes - #1526
Conversation
iamtoruk
left a comment
There was a problem hiding this comment.
Thanks, the design is solid: it reuses src/sqlite.ts read-only, handles the event_zstd column and a missing table, and a session in both the legacy JSONL and the store counts once. Tests pass here.
Three fixes before merge:
- Id-less dedup can drop real calls. In
consumeEvent,dedupId = entry.id ?? h:<hash>. Two id-less assistant events with identical content now collapse into one. Reproduced on a fixture: main 4 calls / $0.010315, this PR 3 / $0.010210 (happens when the event has no timestamp and falls back to the session one). Fix: keep a per-sessionMap<string, number>and useh:${hash}:${n}, withncounting repeats from 0. Event order is the same in both formats, so keys still match across them. - A locked DB during discovery reads as "no sessions". In
discoverSqliteSessions, the catches aroundopenDatabaseand the GROUP BY query turn SQLITE_BUSY into{sources: []}, so that run looks complete but empty and the legacy JSONL of migrated sessions comes back. Fix:if (isSqliteBusyError(err)) throw errin both catches, same as hermes.ts. - The hash change alters cached dedup keys without a parse version bump. Bump
openclaw: 'reported-cost-v1'to'reported-cost-v1-sqlite-store-v1'insrc/session-cache.ts.
Rebase note: main rewrote README.md and dropped the provider table, so keep main's README.
…-01 migration writes OpenClaw 2026.8.1 moved sessions from agents/<agent>/sessions/*.jsonl to agents/<agent>/agent/openclaw-agent.sqlite, so discovery resolved to zero sources and doctor reported the provider empty (getagentseal#1259). transcript_events keeps each JSONL envelope verbatim in event_json, so the JSONL event reducer is extracted and shared: both eras parse, dedup and cost identically, and migrated sessions keep their original ids. Sessions are listed per store and parsed in keyset batches of 2000 rows so a gigabyte store never loads whole. agent-schema 23 lets a row carry its event as an event_zstd BLOB instead; those decode on Node 22.15+ and skip with one notice below that. Fixes getagentseal#1259
7c092f3 to
6982c6f
Compare
|
All three applied, plus the rebase note — kept main's README, so the provider-table row edit went with the table. Id-less dedup: the reducer keeps a per-session Busy discovery: both catches in Parse version:
|
…events before the created_at fallback
Fixes #1259.
Problem
OpenClaw
2026.8.1(2026-09-01) migrated session storage from per-agent JSONL (agents/<agent>/sessions/*.jsonl, archived tosession-sqlite-import-archive/or renamed*.jsonl.deleted.<ts>) to the per-agent SQLite storeagents/<agent>/agent/openclaw-agent.sqlite. The provider still globbed only the legacy path, so every migrated install reported zero sessions —codeburn doctorshowedopenclawasemptyon an actively used install.What changed
src/providers/openclaw.tsnow discovers and reads both storage eras side by side:agent/openclaw-agent.sqlitewith atranscript_eventstable becomes one source per session (<db>:<sessionId>, the forge.ts path convention). Sessions are listed with aGROUP BY session_id; each session parses in keyset batches (seq > ?, LIMIT 2000) so a gigabyte store never loads whole — the reporter's corpus is GB-scale, and overview -p today crashes with heap OOM on large corpora (cache-load layer) #1504/overview -p today still OOMs at 512MB after cache-load bounds (parse retention) #1505 put that class of regression on the record.event_json, so the JSONL parser's state machine (session/model_change/model-snapshot/message) is extracted and both parsers run it — identical parse, cost and dedup semantics across eras, and migrated sessions keep their original session ids (no double-count of history already cached underopenclaw:<sessionId>:<dedupId>).event_zstdBLOB instead ofevent_json(one zstd frame, both sides CHECK-bounded at 4 MiB). The parser detects the column viaPRAGMA table_info, decodes withnode:zlib'szstdDecompressSync(Node 22.15+, the same optional-export treatment as dsh.ts), and skips such rows with a single stderr notice on older runtimes rather than crashing.timestamp→ rowcreated_at(ms, verified against the upstream writer'sDate.now()fallback) → store mtime, mirroring the JSONL file-mtime policy.A store without
transcript_events(not an agent-schema DB) is skipped silently; legacy.clawdbot/.moltbot/.moldbotroots keep working through both eras.Schema evidence
From
openclaw/openclaw:transcript_events(session_id, seq, event_json, created_at)withPRIMARY KEY (session_id, seq)is identical across the shipped schema fixtures v14–v22 (test/fixtures/sqlite/openclaw-agent-schema-v*.sql);event_jsonis the original event envelope (session-accessor.sqlite-transcript-store.tsstringifies the event); the zstd column pair and the 4 MiB CHECKs are the in-flight schema 23 (src/state/openclaw-agent-schema.sql,docs/reference/database-schemas/layout.md: "selectingevent_jsonalone omits compressed events").Tests
11 new cases in
tests/providers/openclaw.test.tsbuilding real stores withnode:sqlite: discovery when the JSONL is gone, parse parity with the JSONL semantics (model, tokens,cost.total, tools/bash), dedup across re-parses,created_atfallback, store-mtime fallback, a 2005-row session across the batch boundary, a DB withouttranscript_events, JSONL+SQLite coexistence, and the zstd column pair (self-skipping below Node 22.15).npx tsc --noEmit— cleannpx vitest run tests/providers/— 51 files, 927/927npx vitest run— 5393 passing; the only failures are the known load-flakytests/cache-refresh-lock.test.ts(passes in isolation) and a worktreeapp/node_modulesartifact, both unrelated and both green on their ownAlso reviewed adversarially by an independent GLM-5.3-Flash pass before submission; its findings that changed the code:
created_atretry on a garbage envelope timestamp, so a present-but-unparseabletimestampno longer lands the call on the store's mtime — "now" on a live gateway.One finding was declined as a deliberate trade-off: discovery opens each agent store per run and lists sessions with one
GROUP BY session_idover the PK index — same shape as forge/cursor-agent discovery, no memory blow-up (the listing is distinct ids, not events); parse-side batching is where the GB-scale guarantee lives.Docs
docs/providers/openclaw.mdand the README provider table row now describe the SQLite store, the two eras and the zstd caveat.