Skip to content

The a540 GPU faults under Skia-GPU are Skia blur/downsample passes stalling SP/TPL1 -- not binning, not fp16, and a different class from the compositor's one VSC fault

scope: soc:msm8998 · severity: finding · confidence: proven · subsystem: gpu

How to read one – /sys/class/devcoredump/devcd*/data (root) right after the fault; it expires. kernel: msm text with the ring and the DUMP-flagged BOs (freedreno marks every cmdstream BO, fd_bo_new_ring). mesa 26.1.6’s crashdec mis-parses this kernel’s revision: 540 (05040001) line (patched locally in ref/mesa-26.1.6) and walks the ring tail, not the hung submit; the reliable route is CP_IB1_BASE (register byte offset 0x2c7c) -> that BO -> repack as .rd (ascii85 words are big-endian; RD_*ADDR records are {lo, len, hi}) -> cffdump -v. The decode names the render target, the sampler/texture descriptors and disassembles the shaders.

What both hung submits are – Skia blur passes: dozens of samb (sample with bias) per pixel walking a downsampled chain, one sampler with MAX_LOD = 0.125 on a texture with MIPLVLS = 0, PIXLODENABLE set in SP_FS_CTRL_REG0. The status says the shader core is waiting on the texture pipe and never retires; the CP is parked. No SMMU stall (the driver checks before printing), no CP_HW_FAULT.

What was excluded, with counts – see evidence line. The one thing that correlates with every decoded hang is the blur shader; the page-without-blur arm still faults because box-shadow is a blur too.

Where it goes – upstream freedreno with the two dumps. Until then Skia-GPU stays off (WEBKIT_SKIA_ENABLE_CPU_RENDERING=1), which is also why a-launch-that-skips-the-user-manager-loses-environment-d matters. The compositor-thread fault (VSC/VPC busy) is a second, unexplained class seen once, under GALLIUM_HUD overlay drawing.