Skip to content

perf: reuse snapshot traversal within element response batches - #1285

Closed
KazuCocoa wants to merge 1 commit into
masterfrom
codex/perf-request-snapshot-context
Closed

KazuCocoa wants to merge 1 commit into
masterfrom
codex/perf-request-snapshot-context

Conversation

@KazuCocoa

@KazuCocoa KazuCocoa commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

This adds a response-local cache for flattened snapshot roots. Live WDA measurements below did not demonstrate a performance benefit: the tested paths either bypassed the context or had no cached root to reuse.

Building a multi-element response can repeatedly flatten the same cached snapshot root while resolving individual queries. Add a context shared only for that response, keyed by root object identity, to reuse the flattened snapshot set.

Continue applying each query's transformers and existing snapshot fallbacks independently. A new response creates a new context, so snapshots are not reused across commands. Tests check one flatten per root, fresh work for new roots or contexts, and release of retained roots when the context ends.

Validation: all 178 native unit tests passed on an iOS 26.5 simulator using a local integration branch containing this review series. The regression tests are included in this PR. Native changes were built with Xcode 27.1.

Live WDA experiment (2026-09-28)

No end-to-end speedup was demonstrated in the tested workloads. The unit tests establish reuse when a shared root is supplied; the live diagnostic found no reuse in these actual element-response paths. This is not currently evidence for merging this as a performance improvement.

Compared baseline 261c08d53917b138eaf2f7d8a744d1151fe3bc8d against this PR alone at 37f1f3e67139a460da6b550aad06289434a1c5b5. Both runners were built in Release with Xcode 27.1 (27A9269) and run on the same iOS 26.5 (23F77) simulator. Both used the same IntegrationApp binary containing a static UIKit grid of 64 or 256 accessible labels, with unique node_i identifiers and Node i & value labels.

Measured real POST /session/{sessionId}/elements HTTP latency, including receiving the complete response body. JSON parsing, assertions and artifact writing were outside the timer. Timed requests used compact responses. Each case had two fresh-session blocks per version, one warmup and five measured requests per block: 10 measured requests per version per case, 100 measured requests total. The first four cases ran before → after → after → before; the added non-native-cache case ran after → before → before → after. Settings and capabilities matched between versions for each case.

Case Before median (ms) After median (ms) Latency reduction
Class name, 64 elements 883.34 894.07 -1.21%
Class name, 256 elements 7534.02 7487.03 +0.62%
Predicate, 64 elements 1198.72 1206.82 -0.68%
Class name, 256 elements, bound by index 6523.82 6498.07 +0.39%
Class name, 64 elements, bound by index, non-native caching 1962.71 1967.06 -0.22%

Positive reduction means faster; negative means slower. These small changes, spanning −1.21% to +0.62%, do not demonstrate a reliable performance benefit. The small sample count also does not establish a small regression or strict performance equivalence.

Class-name lookup used XCUIElementTypeStaticText. The predicate was type == 'XCUIElementTypeStaticText' AND name BEGINSWITH 'node_'. Native caching and accessibility-element binding were enabled by default. The fourth case set boundElementsByIndex=true; the fifth additionally set the session capability useNativeCachingStrategy=false.

Result checks

All 100 measured requests and 20 warmups returned the expected count and ordered element IDs, matching the expanded response collected in the same session. Before and after every timed block, a separate expanded response checked type, label, rect, attribute/name, attribute/index and enabled (40 expanded responses total). Across versions and sessions, these expanded responses were identical after removing only the two element-ID fields. Within each session, the entire expanded element array, including IDs, matched before and after timing. This supports correctness for the tested fixtures; it does not establish universal regression freedom.

Why there was no measurable benefit

After completing all timing runs, a separate local diagnostic build added counters for context lookups, absent roots, flatten operations and cache hits. This instrumentation was excluded from the timing comparison and is not part of the PR.

Each diagnostic case returned 64 elements and ran expanded → compact → expanded requests:

Diagnostic case Lookups per response Missing roots Flatten operations Cache hits
Native caching, default binding 0 0 0 0
Native caching, index binding 0 0 0 0
Non-native caching, index binding 64 64 0 0

All three responses in each case produced the same counters. In the native-cache path, element registration calls fb_cacheId → fb_uid → fb_standardSnapshot, which sets lastSnapshot before the response dictionary is built. The existing lastSnapshot branch therefore bypasses the new context. With non-native caching, the context was consulted, but each query's rootElementSnapshot was nil, so it fell back without flattening or reusing anything.

Recommendation: keep this deferred until a real workload demonstrates repeated use of the same non-nil root in this response path. A synthetic shared-root unit test alone is insufficient to justify the optimization.

Scope and reproducibility

This measured findElements, the path changed by this PR, rather than /source. It covers a static UIKit fixture on one simulator and SDK, not physical devices, changing hierarchies, XPath, or every query shape. The initial pilot and an interrupted 1,024-element validation probe were excluded; the final matrix was limited to 64/256 elements because baseline 256-element requests already took roughly 6–8 seconds.

For each revision, build WebDriverAgentRunner with build-for-testing, -configuration Release, CODE_SIGNING_ALLOWED=NO, USE_PORT=8213, and USE_IP=127.0.0.1; launch with test-without-building, -only-testing:WebDriverAgentRunner/UITestingUITests/testRunner, and -parallel-testing-enabled NO. Launch the identical fixture with the requested label count, create a fresh WDA session per case, apply the settings above, and issue the matching locator requests with one warmup followed by five samples. Restart the runner in the balanced order above. Raw samples, response captures, fixture patch, runner script, summary script and diagnostic patch were retained as local experiment artifacts.

Follow-up: 100 elements sharing an accessibility ID (2026-09-28, Pacific time)

No effective real-world path was found in this follow-up. A live UIKit fixture contains 100 distinct accessible UILabels, all with accessibilityIdentifier = @"something". Labels remain unique (Node i & value) to validate identity and ordering. Real findElements("accessibility id", "something") calls returned all 100 elements. Eight diagnostic scenarios produced zero response-context cache hits.

Diagnostic exploration

The PR head 37f1f3e67139a460da6b550aad06289434a1c5b5 was built with temporary counters and run as an actual WDA XCTest runner on an iOS 26.5 simulator with Xcode 27.1. For each scenario, a fresh app session made expanded → compact → expanded element requests. The counters below were identical for all three responses in each scenario:

Diagnostic case Lookups Missing roots Flatten operations Cache hits
Accessibility ID, default caching/binding 0 0 0 0
Accessibility ID, native caching + index binding 0 0 0 0
Accessibility ID, non-native caching + default binding 100 100 0 0
Accessibility ID, non-native caching + index binding 100 100 0 0
XPath, non-native caching + index binding 100 100 0 0
Class chain, non-native caching + index binding 100 100 0 0
Parent-scoped accessibility ID, non-native caching + index binding 100 100 0 0
Accessibility ID after /source, non-native caching + index binding 100 100 0 0

non-native caching means the session capability useNativeCachingStrategy=false; index binding means boundElementsByIndex=true. The XPath was //XCUIElementTypeStaticText[@name='something']; the class chain was **/XCUIElementTypeStaticText[`name == 'something'`]. Parent-scoped search first found the containing view by its unique SourceBenchmark identifier, then used /session/{sid}/element/{parentId}/elements. The final scenario fetched /source immediately before the first element search in that session.

The parent-scoped scenario also produced three separate one-use contexts, each with one flatten and zero hits, from calls without a shared response context. These are reported separately in the artifacts; they are not response-cache reuse. Although a root was available elsewhere in the lookup path, the 100 returned elements still had missing query roots at the point where this PR tries to reuse them.

These tests do not manually assign rootElementSnapshot, inject synthetic query roots, or change production lookup behavior to force a cache hit. The only WDA changes in this diagnostic pass were counters.

Uninstrumented before/after timing

After diagnosis, WDA sources were restored to the exact PR head and rebuilt without counters. Compared against baseline 261c08d53917b138eaf2f7d8a744d1151fe3bc8d using the same separately-built shared-ID fixture binary, simulator and settings. Release builds; HTTP POST /session/{sid}/elements, compact responses; full response-body transfer included, JSON validation and file writes outside the timer.

Two representative accessibility-ID cases ran in before → after → after → before order, with fresh sessions per case. Each block had one compact warmup and five measured requests: 10 measurements per version per case, 40 measured requests total. Positive reduction means faster; negative means slower.

Case Before median (ms) After median (ms) Latency reduction
Accessibility ID, default caching/binding 2488.26 2528.36 -1.61%
Accessibility ID, non-native caching + index binding 4890.16 5050.10 -3.27%

Per-block medians (ms):

Case Before 1 After 1 After 2 Before 2
Accessibility ID, default caching/binding 2414.10 2540.38 2462.51 2535.35
Accessibility ID, non-native caching + index binding 4875.20 5062.09 5038.77 4903.24

The small timing sample is descriptive; it does not establish a small regression or strict equivalence. More directly, the diagnostic found no reuse mechanism active in these scenarios, so this follow-up supplies no evidence for a speedup attributable to this cache.

Correctness and scope

All timed requests and eight warmups returned 100 unique element IDs in the same order as the expanded response for their session. Sixteen expanded validation responses checked type, label, rect, attribute/name, attribute/index and enabled. Across before/after versions, each case's expanded responses matched after removing only the two element-ID fields. Within every session, complete expanded responses including IDs matched before and after timing. The eight diagnostic scenarios also returned the expected 100 distinct elements with ordered labels, with unchanged expanded responses within each session.

Only the two accessibility-ID cases received balanced before/after timing; the other six cases were diagnostic exploration of the PR head, not performance comparisons. This is a static fixture on one simulator/SDK, not proof that no other XCTest version or lookup path can benefit. A useful real case remains unconfirmed, so the recommendation to defer this PR is unchanged.

The local experiment archive retains the exact fixture patch, HTTP driver, settings/capabilities, raw responses, measurements, build fingerprints, diagnostic patch and logs. No fixture or diagnostic changes were committed or pushed.

@KazuCocoa

Copy link
Copy Markdown
Member Author

I couldn't find beneficial cases with this. Closing

@KazuCocoa KazuCocoa closed this Sep 29, 2026
@KazuCocoa
KazuCocoa deleted the codex/perf-request-snapshot-context branch September 29, 2026 06:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant