GPU rasterisation is visibly wrong on a540 because the a5xx GMEM store never resolves multisample buffers
scope: soc:msm8998 · severity: finding · confidence: proven · subsystem: gpu
The question — with WEBKIT_SKIA_ENABLE_CPU_RENDERING=1 removed, so Skia
rasterises on the GPU, WebKit renders icons and glyphs with blocks of garbage
in them. GPU rasterisation is worth having (it cuts the worst frame stall from
~2000 ms to ~720 ms, see
the-browser-stutter-is-a-blocked-webkit-main-thread), so what is wrong?
The answer — fd5_emit_tile_fini() does not flush the colour cache, and
the GMEM resolve writes through it.
fd5_emit_sysmem_fini(): fd5_emit_tile_fini(): /* GMEM */ fd5_emit_lrz_flush() fd5_emit_lrz_flush() PC_CCU_FLUSH_COLOR_TS fd5_cache_flush() /* UCHE only */ PC_CCU_FLUSH_DEPTH_TS fd5_set_render_mode(BYPASS)fd5_cache_flush() writes UCHE_CACHE_INVALIDATE – the texture/unified
cache. It says nothing about the CCU, and fd5_emit_blit() resolves GMEM with
a CCU_RESOLVE event. A GMEM batch can therefore end with resolved colour
still in the colour cache, and a later batch that samples that texture reads
what was there before. WebKit’s compositor does exactly that, hundreds of times
per frame: render a layer into a small FBO, then sample it.
The measurement. A local page (no network, no video, no YouTube): 80 cards of 5 filled SVG paths = 400 composited layers, scrolled. The CPU rasteriser renders it correctly, so it is the reference; every arm is diffed against it.
| arm | differing pixels |
|---|---|
FD_MESA_DEBUG=sysmem |
0 |
| default (GMEM) | 3604 |
| default, repeat | 4218 |
nobin |
3693 |
flush |
7645, 10070 |
| tile 1024x1024 | 3951 |
| tile 512x512 (single bin) | 2954 |
sysmem is the only correct mode, and the only one that flushes the CCU.
What this rules out, each measured:
- Binning –
nobinis unchanged (3693 vs 3604). - Bin alignment / the trailing partial bin – a 512x512 layer tile is one GMEM bin with no tiling at all and still differs by 2954 px.
- Tile size – 1440, 1024 and 512 all corrupt.
- UBWC – impossible on a5xx anyway:
ubwc_ok = is_a6xx(screen)in freedreno_resource.c and a5xx registers no modifieris_format_supportedhook, so a540 advertises linear only. a5xx cannot render to UBWC either –fd5_gmem.cemit_mrt() hardcodesRB_MRT_FLAG_BUFFERto zero. - Texture tiling –
notiledoes not help. - Blur / backdrop-filter – removing it does not help (4275 vs 3611); the
amplifier is composited layers (985 without
will-change/contain, ~4x less). - Shader/ir3 – the hung-shader theory from the devcoredumps was a count of
sambacross a whole dump, not one shader; the disassembly around the fault is 67 instructions with one properly(sy)-synced sample.
FD_MESA_DEBUG=flush making it worse is consistent: it produces more
batches, and the defect is once per GMEM batch.
The mechanism, narrowed to the shape. The defect follows the shape, not the position on screen. Same page, four icon slots:
| page | differing px | where |
|---|---|---|
| thumb, flag, bookmark, triangle | 4614 | only the triangle’s column |
| all four triangles | 9078 | all four columns |
The triangle is the only one of the four with long diagonal edges, i.e. the
only one with many partially covered pixels. emit_gmem2mem_surf() stores
GMEM with the resolve hardcoded off and the condition commented out beside it:
// bool msaa_resolve = pfb->samples > 1;bool msaa_resolve = false;GMEM holds samples interleaved, so storing a multisample buffer without resolving smears rows – which is exactly the horizontal band of coverage seen across the triangle, and exactly why axis-aligned icons survive.
The framebuffers really are multisampled. Proven by the negative: if
samples were 1 the patched builds below would be no-ops, and they are not –
they change the image drastically.
Two patch attempts, both measured, both WRONG (device reverted to stock after each):
| build | change | mixed | all-tri |
|---|---|---|---|
| r1 stock | – | 4614 | 9078 |
| r2 | PC_CCU_FLUSH_COLOR_TS/DEPTH_TS in fd5_emit_tile_fini() |
4237 | – |
| r3 | msaa_resolve = batch->framebuffer.samples > 1 |
32384 | 68265 |
| r4 | r3 + RAS at source samples, DEST forced MSAA_ONE (the split a4xx makes) |
78914 | 68530 |
r2 changed nothing, so the CCU-flush theory is dead. r3 and r4 made it far worse, so the resolve needs state this note does not know.
What is still true: FD_MESA_DEBUG=sysmem is pixel-exact (0 differing px)
and cuts the worst frame stall from ~2000 ms to ~720 ms.
What the next person needs – not another guess. The correct a5xx GMEM
store sequence for a multisample buffer is only knowable from a command-stream
trace of the Qualcomm blob (how freedreno was reverse engineered in the first
place); the vendor kernel (ref/downstream-wahoo, kgsl) confirms only
gmem_size = SZ_1M for a540, which mainline already matches, and contains no
resolve logic because the resolve lives in the proprietary userspace driver.
a6xx is not a usable model either – it resolves implicitly through its blit
event type, with no equivalent flag.
Building mesa to test this is now a solved problem: pmaports/temp/mesa
carries Alpine’s aport (26.1.6) plus a patch, ~20 min per iteration. Two traps
are baked into that aport copy: pmbootstrap does not evaluate case "$CARCH"
blocks, so conditional makedepends must be hoisted to the top level
(porthole#54), and its crossdirect cross-compiles C/C++ but not Rust, so the
nouveau/rusticl Rust components must be dropped or they link as x86-64.
Meanwhile, FD_MESA_DEBUG=sysmem gives correct GPU rasterisation on this
device, measured: worst frame stall 721/716 ms against 1987/1832 ms for the
shipped CPU rasteriser, at 0 differing pixels. It costs the GMEM bandwidth
advantage, so it is a stopgap for the browser, not a system-wide default.
How to re-check it – the reproducer is icons4.html plus a diff against a
CPU-rasterised capture of the same page. Any arm that is correct without
flushing the CCU would overturn this.
