Last of four. By early October the agents had goals (part 2) and something watching them (part 3). What I didn't have was anywhere to look. Status was spread across eleven Telegram threads, a kanban board, and a folder of cron output files. "What is my team doing right now" was a twenty-minute question.
Before we built the screens, we sent the design to a critic model. It found, among other things, that our "Completed" section would have counted abandoned goals as wins.
Muse, Nous Research's assistant app, answers the same question with three tabs, so I took screenshots of a Muse account and wrote down what each card actually contains. The example text here is from those screenshots, not from our system.
Goals: one card per goal with the objective, state, who owns it, and when it last checked in. A detail panel with acceptance criteria, verification, and the latest judge reason. Grouped as active, waiting on you, paused, done.
Feed: one card per thing that happened. Icon, title, body, maybe a link, thumbs up or down, "Discuss", relative time. "Neon refunded every invoice in full" is a Feed card.
Ideas: one card per thing the agent offers without being asked. "I can keep Buster's care on schedule through Sunday." A reason, a "needs from you", Accept / Dismiss / Discuss, a green check once accepted.
That's what we kept. We'd also wanted a seven-day completion chart, a per-goal activity timeline, and a "tracking" state for background goals. The critic found we didn't have the data for the chart or the timeline. "Tracking" had a different problem.
Hermes profiles can run any model, so we keep one called architect-critic-astra on a high-reasoning GPT-6 variant. Its only job is to read a design doc and the code it claims to build on, and say what's wrong. No praise, under 900 words, verify every claim against source or write "not verified."
We sent it the Goals page design. The review came back with four blockers and source references; it had read fourteen files. Full text is in the repo under docs/design/.
"Cleared" is not "completed." When a long Hermes chat gets compressed, the goal is copied to the new session and the old record is marked cleared. You can also clear one by hand. Our "Completed" section would have counted both as wins. The critic found an actual cleared row in a live database to make the point. Fix: only done counts.
The activity history we promised doesn't exist. Hermes overwrites the goal record in place. The judge writes its verdict straight into that record. There's no completion timestamp. Rebuilding a timeline from chat transcripts would miss goals set with /goal and mix in messages from earlier goals in the same session. Fix: show current state and the latest judge result. Drop the timeline until there's a real event log upstream.
Backend routing was wrong. The dashboard routes a plugin's requests through whichever profile is selected at the time. Installing the backend on one profile wouldn't serve the other ten. Fix: install everywhere, cache per profile, show "unavailable" if a backend lacks the plugin.
"Tracking" would hide failures, and "Needs you" claims more than we know. A paused goal can mean the judge timed out or returned junk, not just healthy waiting. blocked can mean out of scope, not "a human must decide." Fix: group by recorded state (Active, Waiting on you, Paused, Done) and compute a separate "may need attention" flag from the recorded reasons. "Waiting on you" only appears when the judge's reason names a user decision.
I was a little annoyed reading it. Every one of the four was right. We now run the critic on every design doc before a card is cut, and the reviewer profile has the same verify-against-source, cite-lines rule.
Three plugins, because the Hermes dashboard allows one tab per plugin.
goals-page is read-only over every profile's state_meta table. It never writes to a live database.