Skip to content

feat(query): cache results for finished past periods - #174

Merged
ErikBjare merged 4 commits into
masterfrom
query-period-cache
Sep 27, 2026
Merged

ErikBjare merged 4 commits into
masterfrom
query-period-cache

Conversation

@ErikBjare

@ErikBjare ErikBjare commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

Part of ActivityWatch/activitywatch#1465 (Year / Custom range / All time in the Activity view).

aw-webui splits long views into one query per day (ActivityWatch/aw-webui#951), and every page load recomputes every day. On a real 4-year database with a large category set, Year takes ~3 min and All time ~10 min, every time. query2 already takes a cache argument but ignored it, and aw-client-js only caches in the open tab.

This caches query2 results for finished past periods, in memory, keyed by (query text, timeperiod).

Design

What is cached: a (query, timeperiod) result, only if the period ended at least 10 minutes ago. Today and the current hour are always computed fresh. The key is a hash of the exact query text and the period normalized to UTC. The category rules are part of the query text, so editing categories naturally misses the cache.

Invalidation (correctness does not assume the past never changes):

  • Queries only read data through the bucket list (find_bucket) and period-bounded reads (query_bucket, query_bucket_eventcount), which return events whose extent overlaps the period.
  • Every write goes through ServerAPI. After each write it records the time range affected and drops cached entries whose period overlaps it:
    • create_events (including imports and aw-sync style backfills): each event's full extent, plus the old extent of any event replaced by ID.
    • heartbeat: the previous last event and the merged/inserted event. replace_last() overwrites the stored last event, which is not always self.last_event (other writes, out-of-order heartbeats), so the stored last event is read (one LIMIT 1 query per merge) and included too.
    • delete_event: the deleted event's extent.
    • bucket create/update/delete/import: clear everything, since they can change what find_bucket resolves to.
  • Writes run under one lock covering the old-range read, the mutation and the invalidation, so two concurrent updates of the same event can't leave a period in between stale. (SQLite serializes writes anyway.)
  • Overlap is inclusive at the boundaries (conservative: a write exactly at midnight invalidates both days).
  • Race: a query that started before an overlapping write finished may have read old data. Each computation records the write generation it started at, and the store is refused if any overlapping write happened since, or if the bounded write log (10k entries) no longer reaches back that far. Invalidation runs after the write (in finally, so failed writes also invalidate).

Storage: in-memory, bounded LRU (10k entries / 128 MB estimated via JSON size). Deliberately not persisted: a restart starts clean, which also covers any change that bypassed the API (editing the SQLite file while the server is stopped). Year on the database below uses ~14 MB.

Opt-out: enabled by default so aw-webui benefits without changes (aw-client-js sends neither name nor a cache flag). Per request: ?cache=false or "cache": false in the body. Globally: query_cache = false in aw-server.toml.

aw-server-rust: has no query cache either (aw-server/src/endpoints/query.rs evaluates every timeperiod). Not changed here; the same design would port directly (its writes also all go through the datastore API).

Benchmark

Real fullDesktopQuery from aw-webui master (38 statements, 37 KB, 442 regex category rules), one request per day over HTTP, against a copy of a 1.9 GB / ~4-year database (aw-core 686389d, peewee 3.17.6):

requests total slowest
Year, cache=false (today's behaviour) 365 182.3 s 3.18 s
Year, cold (fills cache) 365 158.7 s 2.33 s
Year, warm 365 1.1 s 0.01 s
All time, cold (1106 uncached days + the 365 above) 1471 493.1 s 6.41 s
All time, warm 1471 7.7 s 0.02 s

Memory: 14 MB for Year, 47 MB for all 1471 days.

Filling the cache adds no measurable overhead (cold is within noise of cache=false). The cold path is dominated by categorize; speeding that up is separate work in aw-core.

Tests

tests/test_query_cache.py: hit/miss, key normalization (timezone only), overlap vs non-overlap invalidation for heartbeat / insert / delete / replace-by-ID, bucket changes clearing, merge after an out-of-order heartbeat, bulk-write coalescing, current and future periods never cached, the write-during-computation race, write-log overflow, LRU bounds, per-request and config opt-out, REST flag. Added to make test.

Note (unrelated, found while benchmarking): aw-core allows peewee <5, but peewee 4.x returns BucketModel.created as a datetime, which breaks iso8601.parse_date in aw_datastore/storages/peewee.py on existing databases. The lockfiles pin 3.17, so releases are unaffected.

Clients split long views (Year, All time) into one query per day and recompute
all of them on every load. Past days rarely change, so cache query2 results
in memory keyed by (query text, timeperiod), for periods that ended at least
10 minutes ago.

Every write through ServerAPI records the time range it affected (the full
extent of inserted, replaced, merged or deleted events) and drops overlapping
entries; bucket create/update/delete/import clears the cache. A generation
counter keeps a query that raced with an overlapping write from storing a stale
result. In-memory only (bounded LRU, 10k entries / 128 MB), so a restart
starts clean.

Enabled by default (config: query_cache = true). Clients can bypass it per
request with ?cache=false or "cache": false in the body.
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-27T03:11:51.491811Z a9bb4a0 Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@greptile-apps

greptile-apps Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 4/5

[Medium risk] Adds in-memory query result caching to the server.

The PR should address bulk-write invalidation cost before merging.

Findings

  1. P1 Large writes block cache reads ▶
  2. P2 Callers can change cached results ▶

Summary

The PR adds an in-memory cache for queries over finished periods, with write-based invalidation and request- and config-level opt-outs.

  • Large event batches can make invalidation expensive enough to block cache access.
  • Mutable results are shared with callers.

Diagram

%%{init: {'theme': 'neutral'}}%%
flowchart LR
  Q[Historical query] --> C{Cache hit?}
  C -- Yes --> R[Return stored result]
  C -- No --> E[Evaluate query]
  E --> P[Store result if no overlapping write]
  P --> R
  W[Event write] --> I[Check affected ranges against cached periods]
  I --> C
Loading

Reviews (1) · Last reviewed commit: "feat(query): cache results for finished ..."

Comment thread aw_server/query_cache.py
Comment on lines +144 to +148
stale = [
k
for k, (period, _, _) in self._entries.items()
if any(_overlaps(r, period) for r in ranges)
]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Large writes block cache reads When a bulk import or backfill supplies many events, invalidation checks every event range against every cached period while holding the cache lock. With 100,000 events and 10,000 cached periods, that can mean up to a billion overlap checks, substantially slowing the write and blocking cache reads. Group or coalesce the affected ranges before scanning cached entries.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in d9497c8. Affected ranges are now coalesced before scanning (sorted and merged; past 64 disjoint ranges it falls back to one bounding range, so a big import over-invalidates rather than scanning the cache per event). A write call also logs a single generation, so bulk inserts no longer flood the write log. Worst case is now about 64 x entries overlap checks per write. Tests: test_coalesce_bulk_write_ranges, test_bulk_write_is_one_generation.

Comment thread aw_server/query_cache.py
return None
self._entries.move_to_end(key)
self.hits += 1
return entry[1]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Callers can change cached results The cache stores the original result object and returns it by reference. If an in-process caller changes a returned list or dictionary, later calls for the same query and period receive those changes instead of the query result. Copy mutable results when storing or returning them.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Kept as shared, and documented on QueryCache.get in d9497c8. I measured a deepcopy of a real busy-day result (fullDesktopQuery with 442 category rules, 76 KB JSON) at about 6 ms, which would roughly double warm All time (7.7 s to about 15 s for 1471 days). The only caller is the REST handler, which serializes the result without mutating it. If an in-process caller that mutates results shows up, copying at that call site is the cheaper fix.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: a2f244f04b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread aw_server/api.py Outdated
Comment on lines +407 to +408
self.last_event[bucket_id] = heartbeat
self._last_event_unverified.discard(bucket_id)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Keep out-of-order heartbeat replacements marked unverified

When this heartbeat is older than an existing event, insert does not make it the datastore's last event, but this code clears the verification marker and records the older heartbeat in self.last_event. A subsequent matching heartbeat then calls replace_last, replacing the newer stored event while affected contains only the older heartbeat ranges; consequently, a cached period containing the replaced newer event is never invalidated and continues returning stale data. Keep the bucket unverified unless the inserted heartbeat is confirmed to be the stored last event, or invalidate the actual replace_last target.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch, confirmed. Fixed in 34be7ab: dropped the unverified-marker bookkeeping, and every merge now reads the actually stored last event (one LIMIT 1 query) while the cache is enabled, so whatever replace_last() overwrites is always invalidated. Regression test test_merge_after_out_of_order_heartbeat_invalidates_replaced_event fails without the fix.

…results

A bulk import checked every event range against every cached period under the
lock. Merge overlapping ranges first, fall back to one bounding range past 64,
and log one generation per write call instead of one per event.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: d9497c8b2e

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread aw_server/query_cache.py Outdated

@staticmethod
def key(query: str, period: TimeRange) -> str:
normalized = "\n".join(line.strip() for line in query.splitlines()).strip()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve whitespace inside query literals in cache keys

For queries containing multiline string literals, stripping every physical line changes semantically significant whitespace: for example, RETURN = """a\n b"""; and RETURN = """a\nb"""; produce the same key despite returning different strings. Once the first query is cached for a past period, the second can therefore receive its result; hash the exact query text or use syntax-aware normalization instead.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed. Fixed in a9bb4a0: the key now hashes the exact query text (only the timeperiod is normalized, to UTC). Test covers the multiline-literal case.

replace_last() overwrites the datastore's last event, which is not always
self.last_event: an out-of-order heartbeat is inserted without becoming the
stored last event, so the unverified-marker bookkeeping could miss the range
of the event actually replaced. Read the stored last event (LIMIT 1) on every
merge while the cache is enabled instead.
@ErikBjare

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 34be7ab7e2

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread aw_server/api.py
Comment on lines +268 to +272
for e in events:
if e.id is not None:
old = self.db[bucket_id].get_by_id(e.id)
if old:
affected.append(event_range(old))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Serialize replacements before capturing old ranges

When the threaded server handles two event updates for the same ID concurrently, both requests can read the same old range here before either insert executes. If request A moves the event from day 0 to day 1 and a query then caches day 1 before request B moves it to day 2, B invalidates only day 0 and day 2, leaving the cached day-1 result stale indefinitely. Keep the old-range lookup, mutation, and invalidation atomic; the analogous read-before-write paths in delete_event and heartbeat replacement need the same protection.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed. Fixed in a9bb4a0: all write methods (create/update/delete/import bucket, create_events, delete_event, heartbeat) now run under one RLock covering the old-range read, the mutation and the invalidation, so concurrent replacements of the same event serialize. SQLite serializes writes anyway, so throughput is unchanged in practice.

- Two concurrent updates of the same event could both read the same old range
  before either wrote, leaving a period in between cached stale. Writes now
  run under one RLock covering old-range read, mutation and invalidation.
- Per-line whitespace normalization of the key could conflate queries whose
  string literals differ only in whitespace. Hash the exact text.
@ErikBjare

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. What shall we delve into next?

Reviewed commit: a9bb4a0b92

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@TimeToBuildBob

TimeToBuildBob commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

🤖 AI code review

Adds an in-memory query result cache for finished past periods, keyed by query text and UTC-normalized timeperiod, with invalidation on writes through ServerAPI. Introduces QueryCache class, a write lock serializing mutations, cache opt-out via config and REST flag, and a new test file. Wires the cache through main.py, server.py, rest.py, and config.py.

Needs a look — P2 only

Confidence 4/5

1 finding · ⚠️ 1 P2

⚠️ P2 medium — aw_server/rest.py:335

In rest.py, the query body is obtained via request.get_json(). If the body is not valid JSON, get_json() may return None or raise. The code then does query.get('cache') which would fail if query is None. However, the @api.expect(query, validate=True) should validate the body against the schema, which requires 'query' and 'timeperiods'. If the body is invalid, flask-restx will return a 400 before reaching the handler. So query is not None. But the schema does not require 'cache', so query.get('cache') is safe. The issue is that the cache flag in the body is only checked for is False, not for truthy values. If a client sends 'cache': true, it's ignored and the default is used. That's fine. If a client sends 'cache': 'false' (string), it's not False, so cache remains true. This is a bug as I noted earlier. The fix is to parse the body cache field more robustly. This is a real bug for clients that send string booleans.

body_cache = query.get("cache")
        if body_cache is False or (isinstance(body_cache, str) and body_cache.lower() in ("false", "0", "no", "off")):
            cache = False

How this was verified: Checked rest.py: request.get_json() can return None for empty body, and can return a list if JSON array is sent. The code does not validate that query is a dict before calling .get.

4 advisory findings (summary-only, not scored)

These P2 guard, heuristic, trade-off, or documentation claims are retained for judgment without opening review threads.

⚠️ P2 medium — aw_server/api.py:453

In query2, the cache is checked and stored per timeperiod. The cache key is computed from the query text and period. However, the query text is joined from the list using ''.join(query). The REST layer accepts a list of strings, and the join is done without any separator. If the query list is ['RETURN = 1;', 'RETURN = 2;'], the joined string is 'RETURN = 1;RETURN = 2;' which is a valid query. But if the same query is sent as a single string 'RETURN = 1;RETURN = 2;', the key is the same. That's fine. The real issue: the cache key does not include the 'name' parameter, but name is not used in query execution. However, the cache key also does not include the bucket list or any other context. The invalidation clears on bucket changes, so that's covered. But what about queries that depend on the current time? The cache only caches periods that ended at least 10 minutes ago, so the current time is not a factor. What about queries that use random() or other non-deterministic functions? The cache would return stale results. This is a known limitation, but the PR description says 'Past days rarely change', and the cache is opt-out. This is a trade-off, not a bug. No concrete wrong behavior.

Include name in the cache key, e.g., QueryCache.key(f"{name}\n{query}", period) or add name to the key function.

How this was verified: Checked query2 signature: query2.query(name, query, starttime, endtime, self.db) is called with name. The cache key only uses query and period. The REST layer passes name from request.args, which can vary per request.

⚠️ P2 medium — aw_server/query_cache.py:134

In query_cache.py, the put method's race protection has a flaw: when self._gen != started_gen, it checks the write log for overlapping writes. However, if the write log has overflowed and oldest_logged > started_gen + 1, it returns False. But if oldest_logged <= started_gen + 1, it iterates over all writes in the log. The log is a deque with maxlen, so it contains the most recent writes. If a write happened after started_gen but before the log overflowed, it will be in the log. However, if the log overflowed and the oldest logged generation is exactly started_gen+1, it means the log contains all writes from started_gen+1 onward, so it's fine. The logic seems correct. But there is a subtle issue: the generation counter is incremented on every invalidate call, even if the ranges are empty? No, invalidate returns early if ranges is empty. But clear() calls invalidate with (_MIN,_MAX) which is non-empty. So generation increments. The race protection is sound. However, the put method does not check whether the result is stale due to a write that happened before started_gen but after the query started? Actually started_gen is captured before the query runs, so any write before that is already reflected. So no issue. I don't see a distinct bug here.

How this was verified: Traced the logic: if generation changed and no writes logged, it returns False, which is conservative. If writes are logged, it checks overlap. Correct.

⚠️ P2 medium — aw_server/query_cache.py:128

In query_cache.py, the put method computes the size of the result using json.dumps(result, default=str). This can be expensive for large results, but it's done once per store. The size is used for the byte bound. However, the result is stored as-is, and the size is the JSON serialized length. When the result is later retrieved and serialized by the REST layer, the actual JSON may differ if the result contains non-serializable objects that default=str converts to strings. But the cache stores the result object, not the serialized string, so the size is an estimate. This is acceptable.

How this was verified: The size is an estimate, but the cache stores the object, so it's fine.

⚠️ P2 medium — aw_server/query_cache.py:51

In query_cache.py, the event_range function computes the end as start + event.duration. If event.duration is a timedelta, this works. If it's a float (seconds), it would fail. But aw_core.models.Event.duration is a timedelta, so it's fine.

How this was verified: Event.duration is a timedelta in aw_core.

Files changed (8) — the diff as I read it
  • Makefile — Adds tests/test_query_cache.py to the pytest invocation in the test target.
  • aw_server/api.py — Adds QueryCache instance, write lock, _serialized decorator, invalidation calls in write methods, and cache-aware query2.
  • aw_server/config.py — Adds query_cache = true to default config sections for server and server-testing.
  • aw_server/main.py — Reads query_cache setting from config and passes it to _start.
  • aw_server/query_cache.py — New module implementing QueryCache with LRU, generation-based race protection, and range coalescing.
  • aw_server/rest.py — Adds cache query parameter and body field to QueryResource, passing it to query2.
  • aw_server/server.py — Adds query_cache parameter to AWFlask and _start, passing to ServerAPI.
  • tests/test_query_cache.py — New tests for QueryCache unit behavior and ServerAPI integration invalidation.

Reviewed a9bb4a0b924a · openrouter/deepseek/deepseek-v4-flash-0731 · llm engine · 224s · about this reviewer

Maintainer commands

@TimeToBuildBob review (own line) — fresh review · @TimeToBuildBob fix — a worker acts on the findings. Once per comment; 👀 = received.

Comment thread aw_server/rest.py
name = request.args["name"]
query = request.get_json()
cache = request.args.get("cache", "true").lower() not in ("false", "0", "no")
if query.get("cache") is False:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ P2 — In rest.py, the query body is obtained via request.get_json(). If the body is not valid JSON, get_json() may return None or raise. The code then does query.get('cache') which would fail if query is None. However, the @api.expect(query, validate=True) should validate the body against the schema, which requires 'query' and 'timeperiods'. If the body is invalid, flask-restx will return a 400 before reaching the handler. So query is not None. But the schema does not require 'cache', so query.get('cache') is safe. The issue is that the cache flag in the body is only checked for is False, not for truthy values. If a client sends 'cache': true, it's ignored and the default is used. That's fine. If a client sends 'cache': 'false' (string), it's not False, so cache remains true. This is a bug as I noted earlier. The fix is to parse the body cache field more robustly. This is a real bug for clients that send string booleans.

@ErikBjare
ErikBjare merged commit bd56e20 into master Sep 27, 2026
6 checks passed
@ErikBjare

Copy link
Copy Markdown
Member Author

@TimeToBuildBob and the equivalent for aw-server-rust?

@TimeToBuildBob

TimeToBuildBob commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

No equivalent in aw-server-rust yet. The design ports cleanly:

  • Same entry points: all writes (heartbeat, create_events, delete_event, bucket CRUD) go through endpoint handlers in aw-server/src/endpoints/bucket.rs, which would each call state.query_cache.invalidate(...) after mutation.
  • QueryCache would live in ServerState behind Arc<RwLock<QueryCache>> (the Datastore is already a channel to a serialized worker, so no extra write lock is needed at the endpoint layer).
  • The write-generation counter from the Python side is the only real race guard needed: capture gen before calling aw_query::query(), refuse cache store if gen changed and an overlapping write is in the log.
  • LRU bounds via the lru crate; invalidation logic is a direct port.

Filed ActivityWatch/aw-server-rust#763 to track it.

@TimeToBuildBob

Copy link
Copy Markdown
Contributor

Update: the equivalent now exists — ActivityWatch/aw-server-rust#764, CI green (Android/Ubuntu/macOS/Windows/clippy/fmt), with criterion benchmarks on the Rust workload (cold day-query ~3.4 ms vs ~58 ns cache hit on 5,000 events). Same invalidation design as the Python port: only finished periods past a 10-minute margin are cached, and every write endpoint records its affected extent and drops overlapping entries.

ErikBjare pushed a commit to ActivityWatch/aw-server-rust that referenced this pull request Sep 29, 2026
* feat(query): cache results for finished past periods

Port of ActivityWatch/aw-server#174 to aw-server-rust. Clients such as
aw-webui split long views (Year, All time) into one request per day and
recompute every one on each page load; past days rarely change.

Only periods that ended at least 10 minutes ago are cached. Every write
handler in bucket.rs/import.rs records the extent it affected and drops
overlapping entries; bucket create/delete/import clears the cache. A
write-generation counter refuses to store a result if an overlapping
write raced the query.

Opt out per request with ?cache=false, or globally with
query_cache = false in config.toml.

Refs #763
Refs ActivityWatch/aw-server#174

Git-Session-Id: 72a5

* fix(query): address cache correctness, memory, and profiling review findings

- charge the cache key's own bytes against the byte budget, so a large
  request-supplied query text cannot pin memory while its result is tiny
- hold a write lock across read-old-ranges + write + invalidate in the
  event write handlers, closing the concurrent-replacement window that
  could leave an intermediate period cached
- clear the cache after an import even when it fails partway, since
  earlier buckets may already have been written
- apply the configured query_cache opt-out on Android
- serialize a result exactly once: the cache stores the serialized body
  and the endpoint joins those strings into the response
- add a criterion benchmark measuring the avoided work (cold query vs
  cache hit) on a finished day of 5,000 events

Git-Session-Id: 72a5

* chore(deps): record criterion in Cargo.lock for aw-server

The previous commit added criterion as an aw-server dev-dependency and the
benchmark target without regenerating the lockfile, which breaks --locked
builds.

Git-Session-Id: 72a5
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants