Skip to content

A venus firmware assert wedges the video GDSC, and the driver's recovery then retries every 10 ms forever

scope: soc:msm8998 · severity: finding · confidence: proven · subsystem: media

The question — the decoder stops working mid-session and dmesg fills with System error has occurred, recovery failed to init HFI. What state is the machine actually in, and does it come back?

The answer — it does not come back without a reboot, and while it is not coming back it spins.

1. The firmware asserts. Err_Fatal ... vbuffer.c:569 in the SFR (sub system failure reason) region, session id dead. Seen here during a mid-stream resolution change on 1440p60 VP9, which is what YouTube’s adaptive switching does unprompted.

2. The video GDSC will not power on afterwards. Recovery calls venus_boot() -> pm_runtime_get_sync() -> genpd_power_on() -> gdsc_enable() -> gdsc_toggle_logic(), which times out (-ETIMEDOUT). A full rmmod venus_dec venus_enc venus_core; modprobe venus_core hits the same timeout in venus_probe, so the driver cannot re-take the hardware. /dev/video0-5 disappear and v4l2vp9dec vanishes from GStreamer’s registry – everything silently falls back to software decode from that moment on, which is a trap for any measurement running across it.

3. The recovery handler never gives up. venus_sys_error_handler() (drivers/media/platform/qcom/venus/core.c) ends with:

if (failed) {
disable_irq_nosync(core->irq);
dev_warn_ratelimited(core->dev,
"System error has occurred, recovery failed to %s\n",
err_msg);
schedule_delayed_work(&core->work, msecs_to_jiffies(10));
return;
}

No attempt counter, no backoff. Each pass runs pm_runtime_get_sync(), core_deinit(), venus_shutdown(), hfi_reinit(), venus_boot() and hfi_core_resume(), several of which sit out their own timeouts first.

The log lies about the rate. The message is _ratelimited, so it prints about once a second while the work re-runs every 10 ms plus timeout time. Read the interval between those lines as a retry rate and you will conclude the machine is idle; it is not. irq/158-venus was taking measurable CPU in top the whole time.

Fixed by aport patch 0209-media-venus-bound-the-system-error-recovery-retries.patch (pkgrel 29, kernel #30): count consecutive failures, back off 10/20/40/80 ms, give up after five. core->sys_error stays set on give-up, which every HFI entry point in hfi.c already tests, so calls fail cleanly instead of hanging, and nothing waits on sys_err_done. The hardware still needs a reboot – the patch stops the machine burning CPU until you give it one.

What to check before trusting any media measurement:

Terminal window
sudo dmesg | grep -c "System error has occurred" # must be 0
ls /dev/video* # 0-7 present
gst-inspect-1.0 v4l2vp9dec >/dev/null && echo hw decode present

Related: the-sigkill-venus-wedge-was-vp9-bandwidth-starvation – a different, and recoverable, way for venus to look dead.