A count-based poll loop silently caps how long an experiment may take
scope: generic · severity: trap · confidence: proven · subsystem: tooling
Symptom — a long run reports a dead phone:
>> suspending (alarm=+900s, prep=none)>> resultssh: connect to host 172.16.42.1 port 22: Connection timed outtk_device_state then reads ABSENT, fastboot devices is empty, and it looks
exactly like the resume hang you were hunting. It is not. The phone is still
asleep, and it wakes on its own RTC alarm several minutes later.
Cause — the wait was written as a fixed number of probes, not as a deadline:
for _ in $(seq 1 60); do [ "$(S "grep -q 'TRY DONE' $LOG && echo READY" 12)" = READY ] && breakdoneSixty probes at roughly 6 s each (ConnectTimeout=6 on a host that is not
answering) is about 360 s of patience. Every alarm longer than that reports a
failure it never waited for. Nothing in the loop mentions the alarm, so the cap
is invisible at the call site — ph-suspend-cycle.sh 7200 looks like a
supported thing to run.
Fix — derive the deadline from the experiment’s own duration:
DEADLINE=$(tk_deadline_ms $(( A + ${TK_SUSPEND_MARGIN:-180} )))until tk_expired "$DEADLINE"; do [ "$(S "grep -q 'TRY DONE' $LOG && echo READY" 12)" = READY ] && breakdoneThe general rule — a poll loop’s limit must be a function of what it is
waiting for. seq N encodes a duration in units of “however long a failed
probe happens to take”, which changes with the timeout, the network and the
failure mode, and which no reader can convert back into seconds. When the
duration under test is itself the variable being swept — as it is for a
suspend-hang hunt — a fixed cap silently truncates the sweep at the one place
it matters.
Related — poll-never-sleep is about not sleeping instead of polling; this is the other half: poll to a deadline you can name.
