Skip to content

feat(messagequeue): split partition discovery cadence from polling - #528

Merged
behinddwalls merged 1 commit into
mainfrom
preetam/mq-discovery-interval
Aug 6, 2026
Merged

feat(messagequeue): split partition discovery cadence from polling#528
behinddwalls merged 1 commit into
mainfrom
preetam/mq-discovery-interval

Conversation

@behinddwalls

@behinddwalls behinddwalls commented Aug 6, 2026

Copy link
Copy Markdown
Collaborator

Summary

Why?

Partition discovery, lease acquisition, and worker reconciliation all run on the message-poll ticker (PollIntervalMs, 100ms default). Discovery drives topic-wide work every tick — a DISTINCT partition_key scan, an active-subscriber read, self-lease reads, and lease-acquisition probes — whose query volume multiplies with subscribers × topics at 10x/sec, even though its outcome only changes when membership or the partition set changes. Message polling needs 100ms latency; discovery does not.

What?

New PartitionDiscoveryIntervalMs subscription config (default 1s) drives the supervisor's discovery ticker; PollIntervalMs continues to drive per-partition message polling unchanged. This cuts discovery-driven query volume ~10x at the default settings. The accepted trade-off is that a brand-new partition's first message now waits up to the discovery interval (1s) before a worker picks it up; messages on already-owned partitions are unaffected. Integration test configs pin discovery to 100ms so lease-handoff and rebalance convergence assertions stay fast.

Test Plan

  • ✅ Full Docker integration suite (bazel test //test/integration/extension/messagequeue/...) — discovery-latency-sensitive tests (empty-topic wake-up, rebalance convergence, crash recovery) pass with the new cadence.

Issues

@behinddwalls
behinddwalls force-pushed the preetam/mq-discovery-interval branch from 4f616b7 to 7240e79 Compare August 6, 2026 16:01
@behinddwalls
behinddwalls marked this pull request as ready for review August 6, 2026 16:58
@behinddwalls
behinddwalls requested review from a team and sbalabanov as code owners August 6, 2026 16:58
## Summary

### Why?

Partition discovery, lease acquisition, and worker reconciliation all run on the message-poll ticker (`PollIntervalMs`, 100ms default). Discovery drives topic-wide work every tick — a `DISTINCT partition_key` scan, an active-subscriber read, self-lease reads, and lease-acquisition probes — whose query volume multiplies with subscribers × topics at 10x/sec, even though its outcome only changes when membership or the partition set changes. Message polling needs 100ms latency; discovery does not.

### What?

New `PartitionDiscoveryIntervalMs` subscription config (default 1s) drives the supervisor's discovery ticker; `PollIntervalMs` continues to drive per-partition message polling unchanged. This cuts discovery-driven query volume ~10x at the default settings. The accepted trade-off is that a brand-new partition's first message now waits up to the discovery interval (1s) before a worker picks it up; messages on already-owned partitions are unaffected. Integration test configs pin discovery to 100ms so lease-handoff and rebalance convergence assertions stay fast.

## Test Plan

- ✅ Full Docker integration suite (`bazel test //test/integration/extension/messagequeue/...`) — discovery-latency-sensitive tests (empty-topic wake-up, rebalance convergence, crash recovery) pass with the new cadence.
@behinddwalls
behinddwalls force-pushed the preetam/mq-discovery-interval branch from 7240e79 to 66537f6 Compare August 6, 2026 16:58
@behinddwalls
behinddwalls merged commit d081637 into main Aug 6, 2026
3 of 15 checks passed
@behinddwalls
behinddwalls deleted the preetam/mq-discovery-interval branch August 7, 2026 00:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant