Skip to content

[TEST][UR][L0] Disable USM pooling on Xe2+ devices with L0 driver 1.18+ - #23385

Draft
ldorau wants to merge 9 commits into
intel:syclfrom
ldorau:URL0_Disable_USM_pooling_on_Xe2_devices_with_L0_driver_1.18
Draft

ldorau wants to merge 9 commits into
intel:syclfrom
ldorau:URL0_Disable_USM_pooling_on_Xe2_devices_with_L0_driver_1.18

Conversation

@ldorau

@ldorau ldorau commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Use a proxy pool instead of a disjoint pool in both the L0 v1 and v2
adapters for devices that are Xe2 or newer and have an L0 driver of
version 1.18 or newer. Host pools are not associated with any device,
so pooling of host memory is disabled only if all devices in the context
meet this condition.

Fixes: URT-1298
Blocked-by: NEO-17015

@ldorau
ldorau force-pushed the URL0_Disable_USM_pooling_on_Xe2_devices_with_L0_driver_1.18 branch 3 times, most recently from 687f56b to b1ed852 Compare October 8, 2026 15:23
ldorau added 9 commits October 9, 2026 09:51
Use a proxy pool instead of a disjoint pool in both the L0 v1 and v2
adapters for devices that are Xe2 or newer and have an L0 driver of
at least the minimum version defined by
UR_L0_USM_POOLING_DISABLED_MIN_DRIVER_MAJOR and
UR_L0_USM_POOLING_DISABLED_MIN_DRIVER_MINOR in
common/usm_pooling_disabled.hpp (currently 1.17). Host pools are not
associated with any device, so pooling of host memory is disabled only
if all devices in the context meet this condition.
When USM pooling is disabled for Xe2+ devices with at least the minimum
L0 driver version (see isUsmPoolingDisabled()), the adapter uses a UMF
proxy pool instead of a disjoint pool. The proxy pool's
malloc_usable_size implementation always returns
UMF_RESULT_ERROR_NOT_SUPPORTED, which appendUSMFreeExp() treated as
fatal, causing urEnqueueUSMFreeExp to fail with UR_RESULT_ERROR_UNKNOWN.
This mirrors the tolerance pattern already present in usm.cpp (v1) and
v2/usm.cpp for the same proxy-pool quirk.

Also relax enqueue_alloc.cpp's SuccessReuse test: the pointer-reuse
assertion doesn't hold when pooling is disabled and a proxy pool (with
no reuse semantics) is used, so skip it for such devices via a new
isUsmPoolingDisabledForDevice() helper that mirrors the adapter-side
check and uses the UR_L0_USM_POOLING_DISABLED_MIN_DRIVER_MAJOR/MINOR
constants.
TearDown() called urUSMFree(context, deviceMem) before resetData(),
which enqueues a USM fill on deviceMem. This used the pointer after it
had already been freed.

This was masked when USM pooling kept freed memory mapped/valid for
reuse, but is exposed now that pooling is disabled for Xe2+ devices with
at least the minimum L0 driver version (a UMF proxy pool returns memory
immediately on free), causing urEnqueueUSMFill to fail with
UR_RESULT_ERROR_UNKNOWN in TearDown for every test in the
urGraphPopulatedExpTest hierarchy on affected devices (seen in CI as
widespread exp_graph-test failures on Arc B580).

Fix: reset the data while the allocation is still valid, then free it.
Proxy pools (used when USM pooling is disabled, e.g. on Xe2/BMG devices
with at least the minimum L0 driver version) always fail
malloc_usable_size with UMF_RESULT_ERROR_NOT_SUPPORTED and report size
0. Since AllocationStats relies on that call to track allocation sizes,
UR_USM_POOL_INFO_USED_CURRENT_EXP/USED_HIGH_EXP always reported 0 for
proxy-pool-backed allocations, even though memory was actually in use.

Fix by tracking whether a UsmPool is backed by a proxy pool
(IsProxy/isProxy), and for such pools querying the underlying memory
provider's own statistics (umf.provider.by_handle.{}.stats.
allocated_memory / peak_memory) via umfCtlGet, summing them with the
regular AllocationStats-based accounting used for non-proxy pools.
This mirrors the existing approach already used for
getTotalReservedSize()/getPeakReservedSize().

getTotalUsedSize()/getPeakUsedSize() now return ur_result_t (with the
size returned via an output parameter) instead of throwing through
the UR C API boundary, so UMF errors are propagated cleanly to
urUSMPoolGetInfoExp callers instead of risking an uncaught C++
exception crossing an extern "C" entry point.

Also fix USM/usm_leak_check.cpp: direct_usm() only freed the shared
allocation (p1) and relied on pooling to implicitly free the host/
device allocations (p2/p3) at process exit. With USM pooling disabled
on some devices, this now leaks and fails the test. Explicitly free
p2 and p3 as well.
The Level Zero adapter disables USM pooling on Xe2 or newer Intel GPUs
(IP version >= 20.1.0) with at least the minimum L0 driver version
(UR_L0_USM_POOLING_DISABLED_MIN_DRIVER_MAJOR/MINOR). Detect this
condition in lit.cfg.py from sycl-ls output and mark the tests which
verify pooling behavior (usm_device_read_only.cpp, usm_pooling.cpp) as
UNSUPPORTED on such devices. The minimum driver version is mirrored in
lit.cfg.py as L0_USM_POOLING_DISABLED_MIN_DRIVER_MAJOR/MINOR.
The UMF Level Zero provider always passes the relaxed allocation limits
descriptor to zeMemAllocHost, so some newer L0 drivers accept absurd
sizes like SIZE_MAX. With the disjoint pool this was masked by a failing
umfPoolMallocUsableSize() call, but with the proxy pool (USM pooling
disabled on Xe2+ with at least the minimum L0 driver version) the
allocation succeeds and USM/badmalloc.cpp fails.

Reject host allocations larger than maxMemAllocSize unless relaxed
allocation limits are enabled, matching the V1 adapter behavior.
Many Graph tests run with UR_L0_LEAKS_DEBUG=1 but never free their USM
allocations. This was hidden by the UMF disjoint pool, which releases
all of its slabs on destruction. With USM pooling disabled (Xe2+ with at
least the minimum L0 driver version) the proxy pool forwards every
allocation directly to the driver, so unfreed allocations are reported
as leaks.

Add the missing free() calls.
The UMF proxy pool (used when USM pooling is disabled) does not support
umfPoolMallocUsableSize(). In that case urEnqueueUSMFreeExp (V1) freed
the memory synchronously and returned early without signaling the event
or executing the command list, causing DEVICE_LOST errors and hangs.
V2 appendUSMFreeExp inserted the allocation into the async pool with
size 0, so it could never be reused.

Query the allocation size with zeMemGetAddressRange instead when
umfPoolMallocUsableSize() returns UMF_RESULT_ERROR_NOT_SUPPORTED.
When USM pooling is disabled (Xe2 with at least the minimum L0 driver
version), V1 uses UMF proxy pools, so every USM free goes directly to
zeMemFree and the memory is unmapped immediately. With the disjoint
pool, freed memory stayed mapped, so GPU accesses still in flight were
harmless; now they can trigger GPU page faults, engine resets and
DEVICE_LOST cascades.

The V2 adapter already sets the BLOCKING_FREE policy on its UMF Level
Zero provider. Do the same in V1's L0MemoryProvider when pooling is
disabled by calling zeMemFreeExt with
ZE_DRIVER_MEMORY_FREE_POLICY_EXT_FLAG_BLOCKING_FREE.
@ldorau
ldorau force-pushed the URL0_Disable_USM_pooling_on_Xe2_devices_with_L0_driver_1.18 branch from 8653d12 to cb3cca0 Compare October 9, 2026 09:52
@ldorau

ldorau commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor Author

Fixes: URT-1298
Blocked-by: NEO-17015
Waiting-for: NEO-17015

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant