Skip to content

One browser arm cannot resolve anything under ~10% here: the same build gave 29% and 73% frames-over-budget

scope: device:google-taimen · severity: trap · confidence: proven · subsystem: graphics

Do not compare browser builds with one arm each. On m.youtube.com the frames-over-budget metric swings 29% to 73% on the same binary. A single pair of arms will therefore “show” almost any result you like, in either direction, and it is very easy to write the flattering one down.

This was caught the honest way and only just: an r55 arm gave 29% against a baseline’s 68%, which was reported as “frames over budget more than halved”. A repeat of the same build gave 54%, then 58, 62 and 73. The median was 58% and the real effect was about 1.5% on p50 – inside the noise.

Instead: at least three arms per build, compare medians, and quote the range. Budget ~7 minutes per arm.

Two things that make it worse than ordinary noise:

  • The spread is not the same for every build. r53 is tight (p50 34.2-34.4, 65-69%) while r55 is wide (31.5-34.7, 29-73%). A patch whose benefit depends on the page – damage-driven culling, for instance – adds variance rather than shifting the mean, so “high variance” is itself a signal that the code is doing something, and comparing single arms hides exactly that.
  • Back-to-back arms heat up. Even cooling the die to 48 C between runs, the third arm in a row reached 79 C where the first reached 74 C: the chassis and battery stay warm and the die climbs faster. Cool to a tight floor (taimen-thermal-hygiene / tools/ph-thermal.sh cool) and expect later arms in a series to be slower.

And the standing companion to this: a metric moving is not evidence the code moved it. Instrument the patch to prove it executed – a skip counter, a log line – because “the change had no effect” and “the change never ran” are indistinguishable from the outside and want opposite responses.

The scroll arm has the same problem, and here is its number

Section: The scroll arm has the same problem, and here is its number

Measured 2026-09-05 with tools/ph-scrollarm.sh (pure scroll, no video): eleven control arms across five sweeps gave jank frames >33 ms of 32 36 37 38 38 39 41 42 43 43 44 46 – mean 40.1, sd 4.1.

That sd is the whole story:

  • a single arm carries a 95% interval of +/-8 jank
  • a two-arm mean, +/-5.7
  • so two arms a side resolve only a ~30% change; 20% needs four a side and 10% needs sixteen

It caught a real one. Android’s top-app uclamp read as a clean -20% on two arms (35, 34 against a 43, 42 baseline). Four arms through the daemon gave 39 34 36 42 against 37 36 38 43 – nothing – and a direct replication of the identical configuration came back 37, 38. The hint ships disabled.

Every scheduling lever tried that day landed inside this band: uclamp on the browser, on phoc+phosh, on both, and pinning both CPU clusters to their maximum frequency. Treat “inside the band” as the default expectation, not the surprise. Related: taimen-thermal-hygiene – and never run a second load on the phone during a sweep, a perf report left going took the die to 74.5 C and voided an arm.