Who notices when the agent is stuck?

Part 1 had a seven-hour gap in it. The Gemma 4 goal sat paused from 03:35 to 10:52 on Wednesday while the PR it was "waiting on" got approved, merged, broke a production test, and got a follow-up PR, none of which the goal knew about. Nobody told me either, not loudly enough.

I built a supervisor to catch this. On its first day it misread the same paused goal three hours in a row.

Three ways agents stall

Before writing any of this I went through my logs for what "stuck" looked like. There were blank replies, goals disappearing before the work was done, and jobs retrying the same failed fix. A timeout wouldn't tell me which had happened.

The empty reply. You ask for a multi-step thing. The agent sends back a blank message or an "I'll look into that" and then the turn ends and it never looks into anything. I'd asked my own assistant to add reviewer design context to the team setup and got nothing back. Not an error. Nothing. If you don't happen to be reading the chat at that moment, you'll never know.

"Cleared" isn't "done". A session ends and the goal record is gone. That doesn't mean the goal was met. The engineer agent says "pushed" and nobody checks the remote.

The loop. A cron job fires on a red CI check. The agent tries the same fix. CI stays red. Cron fires again. I found one job's output folder with about 260 run files against the same failing check before anybody looked at it. The check was still red.

Muse, which is what we're copying, nudges a goal after about ten calls with no progress and keeps a dated log you can read. That's per goal, per chat. Our problem was different: eleven profiles sharing one kanban board, and something has to watch all of them and decide which one to poke.

What we built

A Python script collects the state. An hourly agent reads it and decides whether to intervene.

The script is 360 lines, standard library only, and it does nothing but read. It opens every profile's session database and the cron job list and reports idle sessions, empty replies after a user request, paused goals and who paused them, missed heartbeats, and what the supervisor itself did last hour. Then a list of flags. No model involved. Runs in under a second.

The agent side is an hourly cron job on my main profile. It reads the script's output plus a page of rules, and it's allowed to do three things: push a session, repair a heartbeat, or report to me. A push is a message sent into an existing session so the agent wakes up with context instead of starting from zero. If nothing warrants action it answers [SILENT] and that's not delivered anywhere.

The pushes are specific. Not "please continue." More like: "You were asked for reviewer design context and sent back an empty reply. You have no goal. Set one: reviewer rules and design sections added to the setup, with a file or commit as proof."

Day one

Oct 8, Pacific. 24 runs, 6 pushes, 0 heartbeat repairs.

Every push got a reply within the hour. Two replies included the goal that should have existed. One was the summary. Replying is not finishing, though: the de-dup goal replied twice and was still open at midnight.