The Hourly Check-In With No Exit
Part 2 of 2 in What Agents Cost
I'm watching it for CI and review events, with an hourly self check-in armed.
A Claude Code session running Fable 5.1 wrote that line about a pull request (PR) it had opened. I'd started the session from my phone, asked it to finish a task, answered a few of its questions, and put the phone down. About 36 hours later I came back to a finished PR, green and mergeable, and no credits left.
I hadn't known the task was done, and I hadn't noticed it was checking hourly. It was watching nothing. Every hour the check-in fired, the session reloaded its entire conversation and read the same passing CI checks. Still green. It found nothing to do, armed the next check-in, and produced zero commits the whole time.
That was a full week of my Fable 5.1 usage limits, spent checking CI status and nothing else. I'm lucky it ran on a subscription rather than API billing, where the same wakes would have kept charging instead of hitting a cap.
The claim I'd defend from this is narrow. An agent's wait loop has to end on a condition the agent can reach by itself. This one ended when the PR was merged or closed, and both of those are a human's action, so the loop had no exit of its own.
An hourly check-in against a shell loop
The session finished its task in about seven minutes, then started checking CI once an hour. Each check-in was a full inference pass over everything the session had said and read so far, and it kept polling until I ran out of credits.
Every wake paid for the whole session starting over. The harness sent the full prompt and conversation back in, and the model read all of it, called tools to fetch the PR's checks, reasoned over the result, wrote a status line, and armed the next timer. None of that changed from one hour to the next.
The fix I wrote has two parts. The first replaces the timer with a shell script: a background loop that polls the workflow run every 30 seconds and exits on the first terminal state. The container does the polling, so the model takes no turns until the loop exits, then wakes once with one line of output to read. The second is a directive to check a green run at most once, then report and stop.
The first time the script ran, it watched a deploy of 49 jobs that went green in about seven minutes. An hourly check-in armed at the same moment would have first looked about 53 minutes after the run finished, reloading the whole conversation to learn that.
Side by side, the two waits differ on every axis I care about:
- The hourly check-in takes one model turn per hour with no upper bound, reloads the full conversation each time, and ends when a human merges.
- The shell loop takes no model turns while it polls, wakes the model exactly once, and ends when the run finishes, which happens whether or not I'm at my desk.
Why nothing flagged it
Read cold, the status line sounds like the behavior I want from an agent. I don't review a sentence like "I'm watching it" for cost, and I doubt many people do.
Silence was the trigger. A healthy PR emits no events, so the timer fires in exactly the case where there's no work, and it keeps firing for as long as that stays true. Still green.
The instruction also lived where I wasn't looking. It came from the
harness's own prompt, which said to keep an hourly self check-in armed
until the PR merged and to re-arm it on every wake. My repo didn't say
that, and neither did the plugin I vendor into it; afterward I grepped
the plugin and found no reference to send_later,
create_trigger, or re-arming anywhere, and its only CI
rule was a blocking watch that returns. The harness does read a
project skill at .claude/skills/steward/SKILL.md before
acting on a CI event, and it treats that file as the authority on how
proactive to be. Nothing had ever written that file, so the default
was simply what ran.
Before calling this a pattern, I read the docs of five other agent products, and two of them, Devin and OpenAI's Codex cloud, offer scheduled runs as a feature you opt into. The others had nothing that re-wakes the agent, and one of those leaves the polling to the calling client. None I found tells the agent to arm a timer by default, so I'm scoping this to the one harness where I watched it happen.
Webhooks do drop events
The strongest case for the timer is real. A webhook can miss a CI failure or a review comment, and a session that trusts events alone could sit on a red build it never heard about. That's why the default exists.
The recovery for a dropped event is still bounded: a blocking wait with a timeout inside the turn that's already running, or the next real event on the PR, which brings the whole state back with it. A missed notification costs one stale status until something else happens. The timer's failure cost a day and a half of wakes and my whole week of limits, and between those two I'll take the stale status every time.
I'm not arguing that agents should stop watching CI, and I'm not arguing against autonomy. A session that notices a red check on its own PR and fixes it without me is what I want. The wake has to be event-driven or bounded inside a turn. Whatever ends the loop has to be something the agent can reach alone. A loop only a human can end works like a subscription: this one billed a wake every hour until I came back. The PR needed one merge, and it sat for 36 hours waiting on me.
What the steward file says now
The fix shipped as that missing file. It bans scheduled check-ins for any PR, CI run, or deploy, and it overrides the harness default by name. A green run gets one verification read per head commit, then the session reports and ends its turn. A red check on a PR the session opened is work with no round limit until it goes green, or until one of the file's other stop conditions applies, and every stop condition ends the turn with a single comment. None of them schedules anything.
Blocking waits stayed legal, because they return:
gh run watch <run-id> --exit-status --interval 30
Before I let an agent start any loop now, I read its instructions for two things: what ends it, and whether the agent gets there without me. If the second answer is no, the loop doesn't get armed. I never recovered the exact count of wakes behind those 36 hours, and with the session gone I don't expect to.