Skip to content

fix(openclaw): read the per-agent sqlite transcript store the 2026-09-01 migration writes - #1526

Merged
iamtoruk merged 2 commits into
getagentseal:mainfrom
ozymandiashh:fix/1259-openclaw-sqlite-transcript
Sep 28, 2026
Merged

iamtoruk merged 2 commits into
getagentseal:mainfrom
ozymandiashh:fix/1259-openclaw-sqlite-transcript

Conversation

@ozymandiashh

Copy link
Copy Markdown
Collaborator

Fixes #1259.

Problem

OpenClaw 2026.8.1 (2026-09-01) migrated session storage from per-agent JSONL (agents/<agent>/sessions/*.jsonl, archived to session-sqlite-import-archive/ or renamed *.jsonl.deleted.<ts>) to the per-agent SQLite store agents/<agent>/agent/openclaw-agent.sqlite. The provider still globbed only the legacy path, so every migrated install reported zero sessions — codeburn doctor showed openclaw as empty on an actively used install.

What changed

src/providers/openclaw.ts now discovers and reads both storage eras side by side:

  • SQLite era. Each agent's agent/openclaw-agent.sqlite with a transcript_events table becomes one source per session (<db>:<sessionId>, the forge.ts path convention). Sessions are listed with a GROUP BY session_id; each session parses in keyset batches (seq > ?, LIMIT 2000) so a gigabyte store never loads whole — the reporter's corpus is GB-scale, and overview -p today crashes with heap OOM on large corpora (cache-load layer) #1504/overview -p today still OOMs at 512MB after cache-load bounds (parse retention) #1505 put that class of regression on the record.
  • Shared event reducer. The migration copied each JSONL line verbatim into event_json, so the JSONL parser's state machine (session / model_change / model-snapshot / message) is extracted and both parsers run it — identical parse, cost and dedup semantics across eras, and migrated sessions keep their original session ids (no double-count of history already cached under openclaw:<sessionId>:<dedupId>).
  • agent-schema 23 forward-compat. The in-flight schema lets a row carry its event as an event_zstd BLOB instead of event_json (one zstd frame, both sides CHECK-bounded at 4 MiB). The parser detects the column via PRAGMA table_info, decodes with node:zlib's zstdDecompressSync (Node 22.15+, the same optional-export treatment as dsh.ts), and skips such rows with a single stderr notice on older runtimes rather than crashing.
  • Timestamps. Envelope timestamp → row created_at (ms, verified against the upstream writer's Date.now() fallback) → store mtime, mirroring the JSONL file-mtime policy.

A store without transcript_events (not an agent-schema DB) is skipped silently; legacy .clawdbot / .moltbot / .moldbot roots keep working through both eras.

Schema evidence

From openclaw/openclaw: transcript_events(session_id, seq, event_json, created_at) with PRIMARY KEY (session_id, seq) is identical across the shipped schema fixtures v14–v22 (test/fixtures/sqlite/openclaw-agent-schema-v*.sql); event_json is the original event envelope (session-accessor.sqlite-transcript-store.ts stringifies the event); the zstd column pair and the 4 MiB CHECKs are the in-flight schema 23 (src/state/openclaw-agent-schema.sql, docs/reference/database-schemas/layout.md: "selecting event_json alone omits compressed events").

Tests

11 new cases in tests/providers/openclaw.test.ts building real stores with node:sqlite: discovery when the JSONL is gone, parse parity with the JSONL semantics (model, tokens, cost.total, tools/bash), dedup across re-parses, created_at fallback, store-mtime fallback, a 2005-row session across the batch boundary, a DB without transcript_events, JSONL+SQLite coexistence, and the zstd column pair (self-skipping below Node 22.15).

  • npx tsc --noEmit — clean
  • npx vitest run tests/providers/ — 51 files, 927/927
  • full npx vitest run — 5393 passing; the only failures are the known load-flaky tests/cache-refresh-lock.test.ts (passes in isolation) and a worktree app/node_modules artifact, both unrelated and both green on their own

Also reviewed adversarially by an independent GLM-5.3-Flash pass before submission; its findings that changed the code:

  • Store authority at discovery. When a session id exists both as a legacy file and in the store, only the store's source is kept. Keeping both live would let the legacy file's cached turns suppress the store's first parse (the parse-time dedup set is seeded from unchanged sources' cached turns), baking an empty result into the store's session-cache entry; once the legacy file is archived and its non-durable entry evicted, the imported history would be gone from every report until the store changed.
  • created_at retry on a garbage envelope timestamp, so a present-but-unparseable timestamp no longer lands the call on the store's mtime — "now" on a live gateway.
  • Payload-hash dedup fallback for id-less envelopes, replacing the parse-position index so distinct id-less calls cannot collide across (or within) eras while true duplicates still collapse.
  • A stderr notice when a compressed row fails to decode, instead of silently understating usage.

One finding was declined as a deliberate trade-off: discovery opens each agent store per run and lists sessions with one GROUP BY session_id over the PK index — same shape as forge/cursor-agent discovery, no memory blow-up (the listing is distinct ids, not events); parse-side batching is where the GB-scale guarantee lives.

Docs

docs/providers/openclaw.md and the README provider table row now describe the SQLite store, the two eras and the zstd caveat.

@ozymandiashh ozymandiashh added area: cli The codeburn CLI and its core parsing/reporting engine provider: openclaw OpenClaw provider bug Something isn't working needs-real-data-proof PR needs screenshot or output proving it works against real data labels Sep 25, 2026

@iamtoruk iamtoruk left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, the design is solid: it reuses src/sqlite.ts read-only, handles the event_zstd column and a missing table, and a session in both the legacy JSONL and the store counts once. Tests pass here.

Three fixes before merge:

  1. Id-less dedup can drop real calls. In consumeEvent, dedupId = entry.id ?? h:<hash>. Two id-less assistant events with identical content now collapse into one. Reproduced on a fixture: main 4 calls / $0.010315, this PR 3 / $0.010210 (happens when the event has no timestamp and falls back to the session one). Fix: keep a per-session Map<string, number> and use h:${hash}:${n}, with n counting repeats from 0. Event order is the same in both formats, so keys still match across them.
  2. A locked DB during discovery reads as "no sessions". In discoverSqliteSessions, the catches around openDatabase and the GROUP BY query turn SQLITE_BUSY into {sources: []}, so that run looks complete but empty and the legacy JSONL of migrated sessions comes back. Fix: if (isSqliteBusyError(err)) throw err in both catches, same as hermes.ts.
  3. The hash change alters cached dedup keys without a parse version bump. Bump openclaw: 'reported-cost-v1' to 'reported-cost-v1-sqlite-store-v1' in src/session-cache.ts.

Rebase note: main rewrote README.md and dropped the provider table, so keep main's README.

…-01 migration writes

OpenClaw 2026.8.1 moved sessions from agents/<agent>/sessions/*.jsonl to
agents/<agent>/agent/openclaw-agent.sqlite, so discovery resolved to zero
sources and doctor reported the provider empty (getagentseal#1259).

transcript_events keeps each JSONL envelope verbatim in event_json, so the
JSONL event reducer is extracted and shared: both eras parse, dedup and
cost identically, and migrated sessions keep their original ids. Sessions
are listed per store and parsed in keyset batches of 2000 rows so a
gigabyte store never loads whole. agent-schema 23 lets a row carry its
event as an event_zstd BLOB instead; those decode on Node 22.15+ and skip
with one notice below that.

Fixes getagentseal#1259
@ozymandiashh
ozymandiashh force-pushed the fix/1259-openclaw-sqlite-transcript branch from 7c092f3 to 6982c6f Compare September 28, 2026 01:16
@ozymandiashh

ozymandiashh commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator Author

All three applied, plus the rebase note — kept main's README, so the provider-table row edit went with the table.

Id-less dedup: the reducer keeps a per-session Map<hash, number> and the key is h:<hash>:<n> counting repeats from 0. Your repro shape is a fixture now (two id-less assistant events, no envelope timestamp, otherwise identical payloads → 2 calls, not 1), and a second test parses the same session from the JSONL and then from the store through one seenKeys set: the store parse adds nothing, since both eras walk the events in the same order and assign the same keys.

Busy discovery: both catches in discoverSqliteSessions rethrow on isSqliteBusyError, same as hermes.ts. A test holds BEGIN EXCLUSIVE across discoverSessions() and expects the busy shape instead of a complete-but-empty run.

Parse version: openclaw bumped to reported-cost-v1-sqlite-store-v1.

npx tsc --noEmit clean; tests/providers/ 51 files, 933/933. Full suite on the rebased branch: 5505 passing, one load-flaky cli-budget case that passes in isolation, plus the known app/renderer jsdom artifact of this worktree's app/node_modules.

@iamtoruk
iamtoruk merged commit b8a9f3c into getagentseal:main Sep 28, 2026
26 of 27 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area: cli The codeburn CLI and its core parsing/reporting engine bug Something isn't working needs-real-data-proof PR needs screenshot or output proving it works against real data provider: openclaw OpenClaw provider

Projects

None yet

Development

Successfully merging this pull request may close these issues.

OpenClaw provider returns empty after 2026-09-01 SQLite session storage migration

2 participants