You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
[GSD-13592] Kernel private surface left unbound after evictUnusedAllocations() triggered by a transient userptr VM_BIND failure (xe hmm EBUSY→ENOMEM under host memory pressure) → Engine memory CAT error on Arc Pro B70 (BMG G31) #1010
Installation method: the .deb packages from the GitHub releases of compute-runtime (26.31.39395.13 / 26.35.39758.10), IGC (2.40.13 / 2.41.5) and gmmlib (22.10.0), installed with dpkg -i into the container image; checksums verified against the release assets
Level Zero loader 1.32.0; oneAPI 2026.1.1 runtime; PyTorch 2.13/2.14 +xpu
Containers get --device /dev/dri and the /dev/dri/by-path mount; debug keys only for diagnosis (NEOReadDebugKeys=1, PrintDebugSettings=1)
No custom kernel parameters for the GPU
Linux Distribution
Other (please specify below)
Other Linux Distribution
Host: Ubuntu 24.04 with mainline kernel 7.2.0-rc5 (xe); user space: Ubuntu 24.04 containers
~1 in 4 warm boots of a 121 GB model (vLLM, tensor parallel 4, PyTorch XPU graphs = UR command buffers, isUpdatable=false, isInOrder=true) dies at graph capture with Engine memory CAT error [18]: class=ccs page faults at one fixed VA per card. NEO allocation logging maps each VA to offset 0 of a 3 MiB PRIVATE_SURFACE (per-kernel; freed by zeKernelDestroy at teardown).
Mechanism (bpftrace on xe tracepoints + kretprobes, stacks symbolised with the NEO dbgsym):
Model loading copies ~800 MiB tensors from pageable host memory. appendMemoryCopy does not stage them
(OSInterface::isSizeWithinThresholdForStaging: < 512 MiB, < 64 MiB if 2 MiB-aligned), so NEO creates a userptr
allocation (allocateGraphicsMemoryForNonSvmHostPtr).
The host is under reclaim/compaction during the load (kcompactd + kswapd ≈ 2.6 M MMU-notifier invalidations per load).
xe's userptr pin (xe_vma_userptr_pin_pages → drm_gpusvm_get_pages → hmm_range_fault) keeps getting -EBUSY (10,197 of
10,944 calls) until its 1 s timeout; xe converts that -EBUSY to -ENOMEM for the ioctl (by design, "memory pressure causing
HMM range fault timeouts").
Drm::bindBufferObject treats the failure as memory exhaustion: evictUnusedAllocations() → unbinds every idle,
non-always-resident, non-locked allocation, including the per-kernel private surface (UNBIND 1–2 ms after the failed bind,
same thread, stack: allocateGraphicsMemoryForNonSvmHostPtr → BufferObject::bind → DrmMemoryOperationsHandlerBind::
evictUnusedAllocations → evictImpl → BufferObject::unbind).
In this workload the surface is never bound again (no later xe_vma_bind of its VA) and the CAT follows. Rank-for-rank:
only ranks that received a post-launch ENOMEM lost the surface and faulted.
(A minimal standalone — bool*bool kernel, captured XPU graph, forced eviction — does re-bind on the next replay/eager
launch, so the missing re-residency depends on something in the full workload we have not isolated.)
Expected Behavior
A transient userptr VM_BIND failure (xe's hmm range-fault timeout under host memory pressure, reported as -ENOMEM) should not leave kernel-internal allocations such as the per-kernel private surface unbound: either no global eviction sweep is triggered by it, or every evicted kernel-internal allocation is made resident again before its kernel next executes. The workload should not fault.
Reproduction Rate
Sporadic ~ 0% - 25%
Steps to Reproduce
Host under reclaim/compaction pressure (a model load close to host RAM size: 121 GB of weights, 123 GB RAM).
Copy large (~800 MiB) tensors from pageable host memory to the device (appendMemoryCopy; above the staging threshold, so NEO pins them as userptr).
Observe xe_vma_userptr_pin_pages -EBUSY until the 1 s timeout -> VM_BIND returns -ENOMEM -> evictUnusedAllocations() unbinds idle allocations including a kernel's private surface.
Launch that kernel (here: during XPU graph capture) -> Engine memory CAT error at the private surface's VA.
Forcing the failure with VmBindWaitUserFenceTimeout = 1 ms reproduces the eviction on a single card; a minimal standalone (bool*bool kernel, captured XPU graph, forced eviction) does re-bind on the next replay/eager launch, so the missing re-residency depends on something in the full workload we have not isolated.
Is this a regression?
Yes, this is a regression - functionality that previously worked is now broken
Pre-submission Checklist
GPU Hardware
4× Intel Arc Pro B70 (BMG-G31, 8086:e223), full 32 GiB BAR, IOMMU off; host AMD Threadripper PRO 3955WX, 123 GB RAM
DRI Devices Information
GPU Detailed Information (lspci output)
sudo lspci -vvv -k -s 0000:03:00.0 (one of the four identical cards)
Driver Version
26.31.39395.13 (IGC 2.40.13); also reproduced on 26.35.39758.10 (IGC 2.41.5)
Installed GPU Driver Packages
The user-space driver runs inside Docker images (Ubuntu 24.04 base); the host provides only the xe kernel driver.
Driver Installation Details
dpkg -iinto the container image; checksums verified against the release assets--device /dev/driand the/dev/dri/by-pathmount; debug keys only for diagnosis (NEOReadDebugKeys=1,PrintDebugSettings=1)Linux Distribution
Other (please specify below)
Other Linux Distribution
Host: Ubuntu 24.04 with mainline kernel 7.2.0-rc5 (xe); user space: Ubuntu 24.04 containers
Kernel Version & Boot Parameters
GuC 70.58.0
Actual Behavior
~1 in 4 warm boots of a 121 GB model (vLLM, tensor parallel 4, PyTorch XPU graphs = UR command buffers,
isUpdatable=false,isInOrder=true) dies at graph capture withEngine memory CAT error [18]: class=ccspage faults at one fixed VA per card. NEO allocation logging maps each VA to offset 0 of a 3 MiBPRIVATE_SURFACE(per-kernel; freed byzeKernelDestroyat teardown).Mechanism (bpftrace on xe tracepoints + kretprobes, stacks symbolised with the NEO dbgsym):
appendMemoryCopydoes not stage them(
OSInterface::isSizeWithinThresholdForStaging: < 512 MiB, < 64 MiB if 2 MiB-aligned), so NEO creates a userptrallocation (
allocateGraphicsMemoryForNonSvmHostPtr).xe's userptr pin (
xe_vma_userptr_pin_pages→drm_gpusvm_get_pages→hmm_range_fault) keeps getting -EBUSY (10,197 of10,944 calls) until its 1 s timeout; xe converts that -EBUSY to -ENOMEM for the ioctl (by design, "memory pressure causing
HMM range fault timeouts").
Drm::bindBufferObjecttreats the failure as memory exhaustion:evictUnusedAllocations()→ unbinds every idle,non-always-resident, non-locked allocation, including the per-kernel private surface (UNBIND 1–2 ms after the failed bind,
same thread, stack: allocateGraphicsMemoryForNonSvmHostPtr → BufferObject::bind → DrmMemoryOperationsHandlerBind::
evictUnusedAllocations → evictImpl → BufferObject::unbind).
only ranks that received a post-launch ENOMEM lost the surface and faulted.
(A minimal standalone — bool*bool kernel, captured XPU graph, forced eviction — does re-bind on the next replay/eager
launch, so the missing re-residency depends on something in the full workload we have not isolated.)
Expected Behavior
A transient userptr VM_BIND failure (xe's hmm range-fault timeout under host memory pressure, reported as -ENOMEM) should not leave kernel-internal allocations such as the per-kernel private surface unbound: either no global eviction sweep is triggered by it, or every evicted kernel-internal allocation is made resident again before its kernel next executes. The workload should not fault.
Reproduction Rate
Sporadic ~ 0% - 25%
Steps to Reproduce
appendMemoryCopy; above the staging threshold, so NEO pins them as userptr).xe_vma_userptr_pin_pages-EBUSY until the 1 s timeout -> VM_BIND returns -ENOMEM ->evictUnusedAllocations()unbinds idle allocations including a kernel's private surface.Engine memory CAT errorat the private surface's VA.Forcing the failure with
VmBindWaitUserFenceTimeout= 1 ms reproduces the eviction on a single card; a minimal standalone (bool*bool kernel, captured XPU graph, forced eviction) does re-bind on the next replay/eager launch, so the missing re-residency depends on something in the full workload we have not isolated.Is this a regression?
Unknown: only 26.31 and 26.35 were tested.
System Logs / dmesg Output
Additional Information
Workaround we ship. Chunk large pageable H2D weight copies (≤ 32 MiB) so NEO stages them: 0 userptr binds, 0 ENOMEM, 0 evictions; also faster
(800 MiB: 11.1 → 15.2 GB/s).
Asks.
allocation.
evictUnusedAllocations()sweep, guarantee that kernel-internal allocations (per-kernel private surfaces) arere-made resident before their kernel next executes — [GSD-13279] Silent wrong results on Arc A770 (DG2): reused per-dispatch private surface never made resident again — evictUnusedAllocations() unbinds it permanently (regression since ee21f7c717) #973 fixed this for the per-dispatch reuse cache only.
Related: #973 (same trigger, per-dispatch surface); xe "Convert return of -EBUSY to -ENOMEM in VM bind
IOCTL".