The a540 GPU faults under Skia-GPU are Skia blur/downsample passes stalling SP/TPL1 -- not binning, not fp16, and a different class from the compositor's one VSC fault
scope: soc:msm8998 · severity: finding · confidence: proven · subsystem: gpu
How to read one – /sys/class/devcoredump/devcd*/data (root) right after
the fault; it expires. kernel: msm text with the ring and the DUMP-flagged
BOs (freedreno marks every cmdstream BO, fd_bo_new_ring). mesa 26.1.6’s
crashdec mis-parses this kernel’s revision: 540 (05040001) line (patched
locally in ref/mesa-26.1.6) and walks the ring tail, not the hung submit;
the reliable route is CP_IB1_BASE (register byte offset 0x2c7c) -> that BO
-> repack as .rd (ascii85 words are big-endian; RD_*ADDR records are
{lo, len, hi}) -> cffdump -v. The decode names the render target, the
sampler/texture descriptors and disassembles the shaders.
What both hung submits are – Skia blur passes: dozens of samb
(sample with bias) per pixel walking a downsampled chain, one sampler with
MAX_LOD = 0.125 on a texture with MIPLVLS = 0, PIXLODENABLE set in
SP_FS_CTRL_REG0. The status says the shader core is waiting on the texture
pipe and never retires; the CP is parked. No SMMU stall (the driver checks
before printing), no CP_HW_FAULT.
What was excluded, with counts – see evidence line. The one thing that
correlates with every decoded hang is the blur shader; the page-without-blur
arm still faults because box-shadow is a blur too.
Where it goes – upstream freedreno with the two dumps. Until then
Skia-GPU stays off (WEBKIT_SKIA_ENABLE_CPU_RENDERING=1), which is also
why a-launch-that-skips-the-user-manager-loses-environment-d matters.
The compositor-thread fault (VSC/VPC busy) is a second, unexplained class
seen once, under GALLIUM_HUD overlay drawing.
