A venus firmware assert wedges the video GDSC, and the driver's recovery then retries every 10 ms forever
scope: soc:msm8998 · severity: finding · confidence: proven · subsystem: media
The question — the decoder stops working mid-session and dmesg fills with
System error has occurred, recovery failed to init HFI. What state is the
machine actually in, and does it come back?
The answer — it does not come back without a reboot, and while it is not coming back it spins.
1. The firmware asserts. Err_Fatal ... vbuffer.c:569 in the SFR (sub
system failure reason) region, session id dead. Seen here during a
mid-stream resolution change on 1440p60 VP9, which is what YouTube’s
adaptive switching does unprompted.
2. The video GDSC will not power on afterwards. Recovery calls
venus_boot() -> pm_runtime_get_sync() -> genpd_power_on() ->
gdsc_enable() -> gdsc_toggle_logic(), which times out (-ETIMEDOUT).
A full rmmod venus_dec venus_enc venus_core; modprobe venus_core hits the
same timeout in venus_probe, so the driver cannot re-take the hardware.
/dev/video0-5 disappear and v4l2vp9dec vanishes from GStreamer’s registry
– everything silently falls back to software decode from that moment on,
which is a trap for any measurement running across it.
3. The recovery handler never gives up. venus_sys_error_handler()
(drivers/media/platform/qcom/venus/core.c) ends with:
if (failed) { disable_irq_nosync(core->irq); dev_warn_ratelimited(core->dev, "System error has occurred, recovery failed to %s\n", err_msg); schedule_delayed_work(&core->work, msecs_to_jiffies(10)); return;}No attempt counter, no backoff. Each pass runs pm_runtime_get_sync(),
core_deinit(), venus_shutdown(), hfi_reinit(), venus_boot() and
hfi_core_resume(), several of which sit out their own timeouts first.
The log lies about the rate. The message is _ratelimited, so it prints
about once a second while the work re-runs every 10 ms plus timeout time. Read
the interval between those lines as a retry rate and you will conclude the
machine is idle; it is not. irq/158-venus was taking measurable CPU in top
the whole time.
Fixed by aport patch 0209-media-venus-bound-the-system-error-recovery-retries.patch
(pkgrel 29, kernel #30): count consecutive failures, back off 10/20/40/80 ms,
give up after five. core->sys_error stays set on give-up, which every HFI
entry point in hfi.c already tests, so calls fail cleanly instead of hanging,
and nothing waits on sys_err_done. The hardware still needs a reboot – the
patch stops the machine burning CPU until you give it one.
What to check before trusting any media measurement:
sudo dmesg | grep -c "System error has occurred" # must be 0ls /dev/video* # 0-7 presentgst-inspect-1.0 v4l2vp9dec >/dev/null && echo hw decode presentRelated: the-sigkill-venus-wedge-was-vp9-bandwidth-starvation – a different, and recoverable, way for venus to look dead.
