Skip to content

[GSD-13592] Kernel private surface left unbound after evictUnusedAllocations() triggered by a transient userptr VM_BIND failure (xe hmm EBUSY→ENOMEM under host memory pressure) → Engine memory CAT error on Arc Pro B70 (BMG G31) #1010

Description

@RagingNoper

Pre-submission Checklist

GPU Hardware

4× Intel Arc Pro B70 (BMG-G31, 8086:e223), full 32 GiB BAR, IOMMU off; host AMD Threadripper PRO 3955WX, 123 GB RAM

DRI Devices Information

0 crw-rw---- 1 root video  226,   0 Aug 26 18:48 /dev/dri/card0
0 crw-rw---- 1 root video  226,   1 Sep 27 03:20 /dev/dri/card1
0 crw-rw---- 1 root video  226,   2 Sep 27 03:20 /dev/dri/card2
0 crw-rw---- 1 root video  226,   3 Sep 27 03:20 /dev/dri/card3
0 crw-rw---- 1 root video  226,   4 Sep 27 03:20 /dev/dri/card4
0 crw-rw---- 1 root render 226, 128 Sep 27 03:20 /dev/dri/renderD128
0 crw-rw---- 1 root render 226, 129 Sep 27 03:20 /dev/dri/renderD129
0 crw-rw---- 1 root render 226, 130 Sep 27 03:20 /dev/dri/renderD130
0 crw-rw---- 1 root render 226, 131 Sep 27 03:20 /dev/dri/renderD131

/dev/dri/by-path:
total 0
0 lrwxrwxrwx 1 root root  8 Sep 27 03:20 pci-0000:03:00.0-card -> ../card4
0 lrwxrwxrwx 1 root root 13 Sep 27 03:20 pci-0000:03:00.0-render -> ../renderD131
0 lrwxrwxrwx 1 root root  8 Sep 27 03:20 pci-0000:23:00.0-card -> ../card3
0 lrwxrwxrwx 1 root root 13 Sep 27 03:20 pci-0000:23:00.0-render -> ../renderD130
0 lrwxrwxrwx 1 root root  8 Sep 27 03:20 pci-0000:43:00.0-card -> ../card1
0 lrwxrwxrwx 1 root root 13 Sep 27 03:20 pci-0000:43:00.0-render -> ../renderD128
0 lrwxrwxrwx 1 root root  8 Sep 27 03:20 pci-0000:47:00.0-card -> ../card2
0 lrwxrwxrwx 1 root root 13 Sep 27 03:20 pci-0000:47:00.0-render -> ../renderD129
0 lrwxrwxrwx 1 root root  8 Aug 26 18:48 pci-0000:65:00.0-card -> ../card0
0 lrwxrwxrwx 1 root root  8 Aug 26 18:48 pci-0000:65:00.0-platform-simple-framebuffer.0-card -> ../card0

total 0
drwxr-xr-x 2 root root 240 Sep 27 03:20 .
drwxr-xr-x 3 root root 240 Sep 27 03:20 ..
lrwxrwxrwx 1 root root   8 Sep 27 03:20 pci-0000:03:00.0-card -> ../card4
lrwxrwxrwx 1 root root  13 Sep 27 03:20 pci-0000:03:00.0-render -> ../renderD131
lrwxrwxrwx 1 root root   8 Sep 27 03:20 pci-0000:23:00.0-card -> ../card3
lrwxrwxrwx 1 root root  13 Sep 27 03:20 pci-0000:23:00.0-render -> ../renderD130
lrwxrwxrwx 1 root root   8 Sep 27 03:20 pci-0000:43:00.0-card -> ../card1
lrwxrwxrwx 1 root root  13 Sep 27 03:20 pci-0000:43:00.0-render -> ../renderD128
lrwxrwxrwx 1 root root   8 Sep 27 03:20 pci-0000:47:00.0-card -> ../card2
lrwxrwxrwx 1 root root  13 Sep 27 03:20 pci-0000:47:00.0-render -> ../renderD129
lrwxrwxrwx 1 root root   8 Aug 26 18:48 pci-0000:65:00.0-card -> ../card0
lrwxrwxrwx 1 root root   8 Aug 26 18:48 pci-0000:65:00.0-platform-simple-framebuffer.0-card -> ../card0

GPU Detailed Information (lspci output)

sudo lspci -vvv -k -s 0000:03:00.0 (one of the four identical cards)
03:00.0 VGA compatible controller: Intel Corporation Battlemage G31 [Intel Graphics] (prog-if 00 [VGA controller])
	Subsystem: Intel Corporation Device 1701
	Control: I/O- Mem+ BusMaster+ SpecCycle- MemWINV- VGASnoop- ParErr- Stepping- SERR- FastB2B- DisINTx-
	Status: Cap+ 66MHz- UDF- FastB2B- ParErr- DEVSEL=fast >TAbort- <TAbort- <MAbort- >SERR- <PERR- INTx-
	Latency: 0, Cache Line Size: 64 bytes
	Interrupt: pin ? routed to IRQ 198
	Region 0: Memory at 407fb000000 (64-bit, prefetchable) [size=16M]
	Region 2: Memory at 40800000000 (64-bit, prefetchable) [size=32G]
	Expansion ROM at f3800000 [disabled] [size=2M]
	Capabilities: [40] Vendor Specific Information: Intel Capabilities v1
		CapA: Peg60Dis- Peg12Dis- Peg11Dis- Peg10Dis- PeLWUDis- DmiWidth=x4
		      EccDis- ForceEccEn- VTdDis- DmiG2Dis- PegG2Dis- DDRMaxSize=Unlimited
		      1NDis- CDDis- DDPCDis- X2APICEn- PDCDis- IGDis- CDID=0 CRID=0
		      DDROCCAP+ OCEn- DDRWrtVrefEn+ DDR3LEn+
		CapB: ImguDis- OCbySSKUCap- OCbySSKUEn- SMTCap- CacheSzCap 0x0
		      SoftBinCap- DDR3MaxFreqWithRef100=Disabled PegG3Dis-
		      PkgTyp- AddGfxEn- AddGfxCap- PegX16Dis- DmiG3Dis- GmmDis-
		      DDR3MaxFreq=2932MHz LPDDR3En-
	Capabilities: [70] Express (v2) Endpoint, IntMsgNum 0
		DevCap:	MaxPayload 256 bytes, PhantFunc 0, Latency L0s unlimited, L1 unlimited
			ExtTag+ AttnBtn- AttnInd- PwrInd- RBE+ FLReset+ SlotPowerLimit 0W TEE-IO-
		DevCtl:	CorrErr+ NonFatalErr+ FatalErr+ UnsupReq-
			RlxdOrd+ ExtTag+ PhantFunc- AuxPwr- NoSnoop+ FLReset-
			MaxPayload 256 bytes, MaxReadReq 512 bytes
		DevSta:	CorrErr- NonFatalErr- FatalErr- UnsupReq- AuxPwr- TransPend-
		LnkCap:	Port #0, Speed 2.5GT/s, Width x1, ASPM L0s L1, Exit Latency L0s <64ns, L1 <1us
			ClockPM- Surprise- LLActRep- BwNot- ASPMOptComp+
		LnkCtl:	ASPM Disabled; RCB 64 bytes, LnkDisable- CommClk-
			ExtSynch- ClockPM- AutWidDis- BWInt- AutBWInt- FltModeDis-
		LnkSta:	Speed 2.5GT/s, Width x1
			TrErr- Train- SlotClk- DLActive- BWMgmt- ABWMgmt-
		DevCap2: Completion Timeout: Range B, TimeoutDis+ NROPrPrP- LTR+
			 10BitTagComp+ 10BitTagReq+ OBFF Not Supported, ExtFmt+ EETLPPrefix-
			 EmergencyPowerReduction Not Supported, EmergencyPowerReductionInit-
			 FRS- TPHComp- ExtTPHComp-
			 AtomicOpsCap: 32bit- 64bit- 128bitCAS-
		DevCtl2: Completion Timeout: 50us to 50ms, TimeoutDis-
			 AtomicOpsCtl: ReqEn-
			 IDOReq- IDOCompl- LTR+ EmergencyPowerReductionReq-
			 10BitTagReq- OBFF Disabled, EETLPPrefixBlk-
		LnkCap2: Supported Link Speeds: 2.5GT/s, Crosslink- Retimer- 2Retimers- DRS-
		LnkCtl2: Target Link Speed: 2.5GT/s, EnterCompliance- SpeedDis-
			 Transmit Margin: Normal Operating Range, EnterModifiedCompliance- ComplianceSOS-
			 Compliance Preset/De-emphasis: -6dB de-emphasis, 0dB preshoot
		LnkSta2: Current De-emphasis Level: -6dB, EqualizationComplete- EqualizationPhase1-
			 EqualizationPhase2- EqualizationPhase3- LinkEqualizationRequest-
			 Retimer- 2Retimers- CrosslinkRes: unsupported, FltMode-
	Capabilities: [ac] MSI: Enable+ Count=1/1 Maskable+ 64bit+
		Address: 00000000fee07000  Data: 0025
		Masking: 00000000  Pending: 00000000
	Capabilities: [d0] Power Management version 3
		Flags: PMEClk- DSI- D1- D2- AuxCurrent=0mA PME(D0+,D1-,D2-,D3hot+,D3cold-)
		Status: D0 NoSoftRst+ PME-Enable- DSel=0 DScale=0 PME-
	Capabilities: [100 v1] Alternative Routing-ID Interpretation (ARI)
		ARICap:	MFVC- ACS-, Next Function: 0
		ARICtl:	MFVC- ACS-, Function Group: 0
	Capabilities: [110 v1] Null
	Capabilities: [200 v1] Address Translation Service (ATS)
		ATSCap:	Invalidate Queue Depth: 00
		ATSCtl:	Enable-, Smallest Translation Unit: 00
	Capabilities: [420 v1] Physical Resizable BAR
		BAR 2: current size: 32GB, supported: 256MB 512MB 1GB 2GB 4GB 8GB 16GB 32GB
	Capabilities: [220 v1] Virtual Resizable BAR
		BAR 2: current size: 8GB, supported: 256MB 512MB 1GB 2GB 4GB 8GB 16GB 32GB
	Capabilities: [320 v1] Single Root I/O Virtualization (SR-IOV)
		IOVCap:	Migration- 10BitTagReq+ IntMsgNum 0
		IOVCtl:	Enable- Migration- Interrupt- MSE- ARIHierarchy+ 10BitTagReq-
		IOVSta:	Migration-
		Initial VFs: 4, Total VFs: 4, Number of VFs: 0, Function Dependency Link: 00
		VF offset: 1, stride: 1, Device ID: e223
		Supported Page Size: 00000553, System Page Size: 00000001
		Region 0: Memory at 00000407fc000000 (64-bit, prefetchable)
		Region 2: Memory at 0000041000000000 (64-bit, prefetchable)
		VF Migration: offset: 00000000, BIR: 0
	Capabilities: [400 v1] Latency Tolerance Reporting
		Max snoop latency: 1048576ns
		Max no snoop latency: 1048576ns
	Kernel driver in use: xe
	Kernel modules: xe

Driver Version

26.31.39395.13 (IGC 2.40.13); also reproduced on 26.35.39758.10 (IGC 2.41.5)

Installed GPU Driver Packages

The user-space driver runs inside Docker images (Ubuntu 24.04 base); the host provides only the xe kernel driver.

# container image (t214w)
intel-igc-core-2 2.40.13
intel-igc-opencl-2 2.40.13
intel-ocloc 26.31.39395.13-0
intel-oneapi-runtime-dpcpp-sycl-opencl-cpu 2026.1.1-325
intel-oneapi-runtime-opencl 2026.1.1-325
intel-opencl-icd 26.31.39395.13-0
libigdgmm12:amd64 22.10.0
libze-dev 1.32.0
libze-intel-gpu1 26.31.39395.13-0
libze1 1.32.0
ocl-icd-libopencl1:amd64 2.3.2-1build1

# container image (neo2635)
intel-igc-core-2 2.41.5
intel-igc-opencl-2 2.41.5
intel-ocloc 26.35.39758.10-0
intel-oneapi-runtime-dpcpp-sycl-opencl-cpu 2026.1.1-325
intel-oneapi-runtime-opencl 2026.1.1-325
intel-opencl-icd 26.35.39758.10-0
libigdgmm12:amd64 22.10.0
libze-dev 1.32.0
libze-intel-gpu1 26.35.39758.10-0
libze1 1.32.0
ocl-icd-libopencl1:amd64 2.3.2-1build1

Driver Installation Details

  • Installation method: the .deb packages from the GitHub releases of compute-runtime (26.31.39395.13 / 26.35.39758.10), IGC (2.40.13 / 2.41.5) and gmmlib (22.10.0), installed with dpkg -i into the container image; checksums verified against the release assets
  • Level Zero loader 1.32.0; oneAPI 2026.1.1 runtime; PyTorch 2.13/2.14 +xpu
  • Containers get --device /dev/dri and the /dev/dri/by-path mount; debug keys only for diagnosis (NEOReadDebugKeys=1, PrintDebugSettings=1)
  • No custom kernel parameters for the GPU

Linux Distribution

Other (please specify below)

Other Linux Distribution

Host: Ubuntu 24.04 with mainline kernel 7.2.0-rc5 (xe); user space: Ubuntu 24.04 containers

Kernel Version & Boot Parameters

7.2.0-070200rc5-generic
BOOT_IMAGE=/vmlinuz-7.2.0-070200rc5-generic root=/dev/mapper/ubuntu--vg-ubuntu--lv ro pci=realloc crashkernel=2G-4G:320M,4G-32G:512M,32G-64G:1024M,64G-128G:2048M,128G-:4096M
xe                   4464640  8
drm_gpusvm_helper      65536  1 xe
intel_vsec             24576  2 pmt_telemetry,xe
gpu_sched              69632  1 xe
drm_gpuvm              57344  1 xe
drm_buddy              12288  1 xe
drm_ttm_helper         20480  1 xe
ttm                   163840  2 drm_ttm_helper,xe
drm_exec               16384  2 drm_gpuvm,xe
drm_suballoc_helper    24576  1 xe
drm_display_helper    323584  1 xe
cec                   106496  2 drm_display_helper,xe
video                  81920  1 xe
i2c_algo_bit           20480  3 igb,ast,xe

GuC 70.58.0

Actual Behavior

~1 in 4 warm boots of a 121 GB model (vLLM, tensor parallel 4, PyTorch XPU graphs = UR command buffers, isUpdatable=false, isInOrder=true) dies at graph capture with Engine memory CAT error [18]: class=ccs page faults at one fixed VA per card. NEO allocation logging maps each VA to offset 0 of a 3 MiB PRIVATE_SURFACE (per-kernel; freed by zeKernelDestroy at teardown).

Mechanism (bpftrace on xe tracepoints + kretprobes, stacks symbolised with the NEO dbgsym):

  1. Model loading copies ~800 MiB tensors from pageable host memory. appendMemoryCopy does not stage them
    (OSInterface::isSizeWithinThresholdForStaging: < 512 MiB, < 64 MiB if 2 MiB-aligned), so NEO creates a userptr
    allocation (allocateGraphicsMemoryForNonSvmHostPtr).
  2. The host is under reclaim/compaction during the load (kcompactd + kswapd ≈ 2.6 M MMU-notifier invalidations per load).
    xe's userptr pin (xe_vma_userptr_pin_pages → drm_gpusvm_get_pages → hmm_range_fault) keeps getting -EBUSY (10,197 of
    10,944 calls) until its 1 s timeout; xe converts that -EBUSY to -ENOMEM for the ioctl (by design, "memory pressure causing
    HMM range fault timeouts").
  3. Drm::bindBufferObject treats the failure as memory exhaustion: evictUnusedAllocations() → unbinds every idle,
    non-always-resident, non-locked allocation, including the per-kernel private surface (UNBIND 1–2 ms after the failed bind,
    same thread, stack: allocateGraphicsMemoryForNonSvmHostPtr → BufferObject::bind → DrmMemoryOperationsHandlerBind::
    evictUnusedAllocations → evictImpl → BufferObject::unbind).
  4. In this workload the surface is never bound again (no later xe_vma_bind of its VA) and the CAT follows. Rank-for-rank:
    only ranks that received a post-launch ENOMEM lost the surface and faulted.
    (A minimal standalone — bool*bool kernel, captured XPU graph, forced eviction — does re-bind on the next replay/eager
    launch, so the missing re-residency depends on something in the full workload we have not isolated.)

Expected Behavior

A transient userptr VM_BIND failure (xe's hmm range-fault timeout under host memory pressure, reported as -ENOMEM) should not leave kernel-internal allocations such as the per-kernel private surface unbound: either no global eviction sweep is triggered by it, or every evicted kernel-internal allocation is made resident again before its kernel next executes. The workload should not fault.

Reproduction Rate

Sporadic ~ 0% - 25%

Steps to Reproduce

  1. Host under reclaim/compaction pressure (a model load close to host RAM size: 121 GB of weights, 123 GB RAM).
  2. Copy large (~800 MiB) tensors from pageable host memory to the device (appendMemoryCopy; above the staging threshold, so NEO pins them as userptr).
  3. Observe xe_vma_userptr_pin_pages -EBUSY until the 1 s timeout -> VM_BIND returns -ENOMEM -> evictUnusedAllocations() unbinds idle allocations including a kernel's private surface.
  4. Launch that kernel (here: during XPU graph capture) -> Engine memory CAT error at the private surface's VA.

Forcing the failure with VmBindWaitUserFenceTimeout = 1 ms reproduces the eviction on a single card; a minimal standalone (bool*bool kernel, captured XPU graph, forced eviction) does re-bind on the next replay/eager launch, so the missing re-residency depends on something in the full workload we have not isolated.

Is this a regression?

  • Yes, this is a regression - functionality that previously worked is now broken

Unknown: only 26.31 and 26.35 were tested.

System Logs / dmesg Output

2026-09-25T00:00:43+00:00 kernel: xe 0000:43:00.0: [drm] Tile0: GT0: Engine memory CAT error [18]: class=ccs, logical_mask: 0x1, guc_id=18
2026-09-25T00:00:43+00:00 kernel: xe 0000:43:00.0: [drm] Tile0: GT0: Engine memory CAT error [18]: class=ccs, logical_mask: 0x1, guc_id=18
2026-09-25T00:00:43+00:00 kernel: xe 0000:43:00.0: [drm] Tile0: GT0: Engine memory CAT error [18]: class=ccs, logical_mask: 0x1, guc_id=18

Additional Information

Workaround we ship. Chunk large pageable H2D weight copies (≤ 32 MiB) so NEO stages them: 0 userptr binds, 0 ENOMEM, 0 evictions; also faster
(800 MiB: 11.1 → 15.2 GB/s).

Asks.

  1. Do not respond to a userptr bind -ENOMEM (a transient hmm timeout) with a global eviction sweep; retry, or fail only that
    allocation.
  2. After any evictUnusedAllocations() sweep, guarantee that kernel-internal allocations (per-kernel private surfaces) are
    re-made resident before their kernel next executes — [GSD-13279] Silent wrong results on Arc A770 (DG2): reused per-dispatch private surface never made resident again — evictUnusedAllocations() unbinds it permanently (regression since ee21f7c717) #973 fixed this for the per-dispatch reuse cache only.
  3. Consider staging large pageable copies in chunks instead of pinning them as a userptr above the 512 MiB threshold.

Related: #973 (same trigger, per-dispatch surface); xe "Convert return of -EBUSY to -ENOMEM in VM bind
IOCTL".

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions