perf: use native predicates for accessibility id lookups - #1286
Merged
Merged
Conversation
KazuCocoa
marked this pull request as ready for review
September 29, 2026 06:59
| @autoreleasepool { | ||
| return [[FBXCElementSnapshotWrapper wdNameWithSnapshot:snapshot] isEqualToString:accessibilityId]; | ||
| NSPredicate *predicate = [NSPredicate predicateWithFormat: | ||
| @"(identifier != nil AND identifier != '' AND identifier == %@) OR " |
There was a problem hiding this comment.
why do we need to compare with nil and empty string?
Member
Author
There was a problem hiding this comment.
Since they caused regressions. I have added tests to cover the behavior
mykola-mokhnach
approved these changes
Sep 30, 2026
github-actions Bot
pushed a commit
that referenced
this pull request
Sep 30, 2026
## [16.13.4](v16.13.3...v16.13.4) (2026-09-30) ### Performance Improvements * use native predicates for accessibility id lookups ([#1286](#1286)) ([2ceb940](2ceb940))
|
🎉 This PR is included in version 16.13.4 🎉 The release is available on: Your semantic-release bot 📦🚀 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Accessibility-ID searches currently use an Objective-C block that calls
wdNameWithSnapshot:. Express the same matching rule as a structured XCTest predicate onidentifierandlabel. The existingfb_snapshotBlockPredicateWithPredicate:helper returns an incoming block predicate unchanged, so this change removes block-based matching; it does not eliminate an additional snapshot-wrapper allocation from this accessibility-ID path.A nonempty identifier still takes precedence over the label. Label fallback applies only when the identifier is absent or empty; an empty name remains unmatchable. Preserve first-match extraction, result ordering, and matching the scope element itself. The WDA
idandnamealiases use this same implementation. Predicate-string, class-chain, class-name, and XPath searches are outside this change.Live WDA measurements
findElementsworkloadMeasured the complete HTTP response for
POST /session/{sid}/elementswith{"using":"accessibility id","value":"something"}against running WDA. This includes lookup and element serialization. These are not predicate-only microbenchmarks.261c08d53917b138eaf2f7d8a744d1151fe3bc8d. After: that revision plus exactly the productionXCUIElement+FBFind.mchange in this PR. No other performance PRs were included.c9366a227b6de19381c96d56a548e9fb53d77dba; the intervening screenshot change does not change this lookup implementation. Repository regression tests run on the PR version.The second baseline block was slower, but both after blocks were faster than both baseline blocks in each case. This is a small simulator experiment, not a universal speedup guarantee. It demonstrates end-to-end latency improvement, not where XCTest executes the predicate or how many AX IPC calls it makes. Other locator strategies and the
id/namealiases were not separately benchmarked.Correctness and regression coverage
All 60 timed requests and 12 warmups returned the expected counts and stable ordered identities within each session. Another 24 expanded-response checks compared
type,label,rect,attribute/name,attribute/index, andenabled: identical within sessions including IDs, and identical across versions after removing only the two element-ID fields. Settings and capabilities matched.A separate before/after live HTTP fixture passed identifier-over-label precedence, nil/empty identifier fallback, empty/missing queries, quotes/Unicode, first-match order, parent-scoped lookup, matching the parent itself, and a button-driven identifier mutation (four old-name matches become three, with one new-name match and stable remaining IDs). Saved semantic outputs matched after removing element IDs.
Validation: 8 targeted integration tests passed, 0 failures, on the PR checkout (Xcode 27.1 / iOS 26.5 simulator).
git diff --checkalso passed.The dedicated
testEmptyIdentifierDoesNotMatchEmptyAccessibilityIdregression test documents why the identifier branch must reject empty strings. It explicitly checks fixtures with an empty identifier and either a nonempty or empty label, then verifies empty searches return no matches for both single/all-result requests and bothuseFirstMatchsettings. All 5 identifier-semantics integration tests passed after adding this coverage. As a mutation check, removing onlyidentifier != ''while retainingidentifier != nilmade the dedicated test fail: 10 matches instead of 0. Restoring the guard made it pass again. The test comments explain the existingwdNamefallback and empty-to-nil behavior. This validation ran on Xcode 27.1 / iOS 26.5 simulator.The PR adds a small opt-in IntegrationApp scene and focused regression tests for those matching rules, scoped lookup under both element-binding modes, and identifier changes. Existing identifier and stable-instance tests are also exercised. Supported older XCTest versions, physical devices, WebViews, and custom accessibility implementations have not been exhaustively tested.
Memory follow-up results
Overall lifetime peaks overlap in the ABBA series (before: 78.44–79.22 MiB; after: 78.52–78.58 MiB). However, the focused 60-request follow-up ended at 78.13 MiB after versus 75.17 MiB before, a 2.95 MiB / 3.93% higher idle footprint in that pair. Thus these measurements do not establish memory neutrality. There is no clear pattern of continuous growth proportional to requests in the sampled traces, but the sample size and duration cannot rule out leaks or a small memory regression. The observed timing of footprint rises varies, including a rise during the optimized follow-up's post-request idle window; this is why its post-idle footprint exceeds its sampled request-only peak. Allocator/runtime behavior and the specific allocations responsible were not profiled, so the difference is not attributed to removal of the autorelease pool alone.
Total: 1,040 measured searches plus 70 warmups, six fresh WDA launches. The 920-request ABBA series and the 120-request focused follow-up are reported separately below. All element-count/order/identity assertions passed. Host: Apple M5 Pro, 24 GiB RAM, macOS 26.6.2. These are WDA runner process measurements on a simulator, not whole-device memory measurements.
Memory measurements and limitations
Memory regression check
Measured the same Release binary pair and identical grid fixture as the latency experiment, with the production diff verified identical to PR commit
86ec5147. Xcode 27.1 / iOS 26.5 simulator. Fresh WDA launches in before → after → after → before order. Each launch uses fresh sessions for all three cases, with five warmups followed by 30 requests for 100 matches, and 100 requests each for one/no match. 920 measured requests plus 60 warmups, all with expected counts and stable ordered IDs. Settings/capabilities were equal across runs.An external sampler records WDA runner physical footprint and RSS with
proc_pid_rusage(RUSAGE_INFO_V4), targeting a 20 ms interval. Idle values are medians over the last two seconds of a five-second wait. Growth means idle after repeated searches minus idle after warmup, with the same session still alive. Negative values indicate reclamation/accounting changes, not negative allocation.Physical footprint, MiB. Each range spans the two independent launches of that variant. Request peaks are sampled maxima per launch.
Per-launch idle footprint (warmup → repeated searches → session deletion), in MiB:
Kernel-maintained lifetime peak footprint, including WDA startup and all three cases:
RSS cross-check (sampled request peaks, MiB):
Scope: sampled per-case peaks can miss brief spikes; the kernel lifetime high-water mark is reported separately and cannot be attributed to a single request. RSS includes shared pages, so physical footprint is the primary comparison. This measures the WDA runner, not aggregate memory across the app, accessibility services, and simulator. Two launches per variant and these fixed-size scenes cannot establish absence of leaks in every workload or during hours-long sessions. A clock-calibration pilot was interrupted and excluded before the final series; no completed measured requests or final-series memory samples were discarded.
Focused follow-up: 60 repeated 100-match searches
The original ABBA runs had a low-footprint baseline launch (~65 MiB after 30 requests), while other launches reached ~75 MiB. This prompted an additional fresh before/after pair, restricted to 100 nodes / 100 matches, with five warmups and 60 measured searches per variant. These 120 requests and 10 warmups are reported separately; the original results were not discarded.
Physical footprint during each consecutive group of six requests (median MiB):
The old implementation also exhibits the approximately 10 MiB rise, so that observation alone does not establish a regression caused by removing the per-evaluation autorelease pool. The experiment does not identify the allocations responsible for the rise.
Reproduce the memory measurements
Memory regression experiment for WDA PR #1286
The source change is byte-for-byte identical to the production diff in PR commit 86ec514. This experiment reuses the earlier Release latency benchmark pair: baseline 261c08d, and that baseline plus the locator optimization only. See provenance.json for hashes. The PR's later unrelated master change and regression-test app scene are not part of these benchmark binaries. Both variants use the identical installed grid fixture from ../fixture.patch.
Run order is before-1, after-1, after-2, before-2. Each block launches a fresh WDA runner and creates a fresh session for each of three cases, in the same order:
Each case pauses 5 seconds after warmup, 5 seconds after measured requests, and 5 seconds after deleting its session. Requests are sequential, with no intentional pacing within each measured phase. Compact responses, native caching, and accessibility binding are held constant. Every request asserts the expected element count and stable ordered element IDs. The complete series has 920 measured requests plus 60 warmups.
sample-memory.c queries macOS proc_pid_rusage(RUSAGE_INFO_V4) for the WDA runner process selected by its listening port. It records resident size (RSS), physical footprint, and the kernel-maintained lifetime peak physical footprint roughly every 20 ms. The sampler is an external process, not instrumentation injected into WDA. mach_absolute_time and Python monotonic_ns share the same clock. An interrupted clock-calibration pilot is preserved in clock-calibration-discarded and excluded from all results.
Per-case peaks are maxima over sampled request execution. They can miss spikes shorter than the sampling interval. The kernel lifetime maximum captures the entire process lifetime, including startup and prior cases; it is reported per block, not misattributed to an individual case. Idle values are medians over the final two seconds of each five-second wait. Post-search growth is post-search idle minus post-warmup idle, while the same session remains alive. Negative deltas can result from allocator/runtime reclamation and memory accounting changes. RSS includes shared resident pages, whereas physical footprint is the main comparison metric. These results cover the WDA runner, not aggregate simulator/device memory or accessibility services.
Reproduction
Use the build and fixture instructions in PR #1286. Update simulator UUID and derived-data paths in run-memory.py. Compile the external host sampler:
Output directories must not already exist. summarize-memory.py writes summary.json from raw memory.csv, request records, and phase timestamps. No memory samples or completed measured requests are removed as outliers.
Focused follow-up
The original ABBA runs showed a roughly 10 MiB difference in the early 100-match phase of one baseline launch. To investigate rather than discard that observation, run-followup.py repeats only the 100-node/100-match case with a fresh runner per variant, five warmups, and 60 measured requests (before then after). This is a separate follow-up, not pooled into the original 30-request ABBA results. It adds 120 measured requests and 10 warmups. All original results are retained.
sample-memory.c
run-memory.py
run-series.py
run-followup.py
summarize-memory.py
provenance.json
{ "before": { "library_sha256": "04e2d8f2c4ce8a9664fc11552534ed5e1589b9eb3ddf5eef86e292a9d02bfde2" }, "after": { "library_sha256": "2441970a2ebdc716b371bd4aa0761c98d56be3905bd1ea0b8e22ab6458e33849" }, "fixture_sha256": "688bcace42aaa5240d8751e7794763fafe594ab77132e622c2adcdfd65f25907", "production_diff_identical_to_pr": true, "benchmark_base": "261c08d53917b138eaf2f7d8a744d1151fe3bc8d", "pr_commit": "86ec5147e307628923dfbfef71c934dcb03a3a13", "note": "Same Release binary pair and app fixture as previous latency experiment; only production locator diff differs. No other performance PRs included. Clock-calibration pilot discarded before measurement.", "host": { "cpu": "Apple M5 Pro", "ram_bytes": 25769803776, "macos": "26.6.2" } }Reproducing the experiment
The collapsed files below contain the exact temporary fixture, driver, summarizer, semantic validator, and raw timing samples. They are experiment artifacts, not production changes. Apply
fixture.patchto a separate checkout of the benchmark baseline, build/install its IntegrationApp once, and use that identical app with both WDA runners. The committed small regression fixture is separate from this performance fixture.Build baseline and changed WDA checkouts independently with
xcodebuild -project WebDriverAgent.xcodeproj -scheme WebDriverAgentRunner -configuration Release -destination 'platform=iOS Simulator,id=<UDID>' -derivedDataPath <before-or-after-directory> build-for-testing CODE_SIGNING_ALLOWED=NO USE_PORT=8213 USE_IP=127.0.0.1. Use only the production source diff for the changed benchmark checkout. Build IntegrationApp from the temporary fixture checkout with the same configuration/destination and install the resulting.appusingxcrun simctl install.Save the scripts together, update their simulator UUID, output directory, and derived-data paths for your machine, then run:
fixture.patch
run-block.py
summarize.py
semantics.py
measurements.csv
metadata.json
{ "baseline": "261c08d53917b138eaf2f7d8a744d1151fe3bc8d", "prototype": "baseline plus prototype.patch; no other performance PRs", "prototype_patch_sha256": "156503ddcf1160a0caec3a9acdcd58e0cff9f9aa47d827bf14b2e5955fc60e0e", "xcode": "27.1 (27A9269)", "runtime": "iOS 26.5 (23F77)", "configuration": "Release", "cases": [ "id-100", "selective-1024-1", "missing-1024-0" ], "order": [ "before-1", "after-1", "after-2", "before-2" ], "samples_per_case_per_block": 5, "warmups_per_case_per_block": 1, "semantic_checks_equal": true, "before_binary_sha256": "04e2d8f2c4ce8a9664fc11552534ed5e1589b9eb3ddf5eef86e292a9d02bfde2", "after_binary_sha256": "2441970a2ebdc716b371bd4aa0761c98d56be3905bd1ea0b8e22ab6458e33849" }