The skin thermal zone storms the PMIC ADC at ~1600 interrupts a second, awake and idle, in the shipped config
scope: device:google-taimen · severity: finding · confidence: proven · subsystem: thermal
~1600 interrupts a second, forever, on an idle phone. pm-adc-tm5 has
fired 2.25 million times since boot and is still going at 1572/s, dragging
~670 VADC conversions a second behind it. This is the shipped configuration,
screen on, nothing running.
It is not a trip crossing. skin-thermal reads 32.1 C and its lowest trip
is 38 C. The PMIC’s own comparator is latched: STATUS_LOW (0x340a) reads
0x10 – channel 4 – and stays 0x10 across seconds, across threshold
rewrites, and across a driver unbind/rebind. The latch lives in the PMIC and
nothing mainline does can clear it.
Mainline’s ISR has no path that could. adc_tm5_isr() reads
ADC_TM5_STATUS_LOW/_HIGH, and for any channel whose bit is set and whose
interrupt is enabled it calls thermal_zone_device_update() – and writes
nothing back. It relies entirely on the thermal core coming back through
set_trips. The gen2 ISR right below it does clear, via the dedicated
*_CLR registers. The vendor’s qpnp_adc_tm.c ISR opens with
qpnp_adc_tm_disable(chip) and then clears LOW_THR_INT_EN for the sensor
that fired, re-arming afterwards. Mainline does neither, so once the
comparator latches, the interrupt free-runs.
And the comparator will not unlatch, because its measurement is wrong. The latch holds even when the threshold is moved to the 44 C code (0x25ee), which is further from 32 C in the direction that should clear it. Whatever the ADC-TM samples on channel 4 sits below any threshold that can be programmed, while the IIO/VADC path on the same input reports a correct 32.1 C. The temperature the thermal core sees is right; the hardware monitor behind it is not converting usefully.
The vendor does not use the monitor for this at all. Our own DT comment records it: “vendor thermal-engine samples bd_therm2 every 2000 ms”. The skin sensor is polled downstream, not wired to a hardware threshold monitor. We implemented it as an adc-tm zone, which is both a behaviour mismatch and the source of this storm.
Three ways out, cheapest first.
- Poll it, as the vendor does. Take
skin-thermaloffpm8998_adc_tmand give it a 2000 ms polling zone over the VADC IIO channel. A DTS change, thebootrung. Removes the storm outright and matches downstream. - Make the ISR self-limiting, vendor-shaped: clear
LOW_THR_INT_ENfor the channel that fired before calling into the thermal core, and re-arm inset_trips. Careful – the core skipsset_tripswhen the trip window is unchanged, so a naive disable loses the trip forever. - Find out why channel 4’s ADC-TM measurement reads below every threshold. The most interesting answer and the most work. The config block matches what mainline intends, so the fault is upstream of it.
Do not read this as a suspend problem. During a 31 s s2idle the same interrupt only manages ~47/s; the storm is an awake-and-idle cost, which is exactly where this port’s 378 mA blanked floor lives.
