The agent sets its own goal. Define done before starting.

Last post I described a feature that eleven agents merged while I was asleep. The agents could finish tasks. They were worse at knowing whether the result I wanted had actually happened.

The tool that fixes some of that is 398 lines including tests. Getting the agent to say what "done" meant before it started took most of a week.

Where this came from

Muse is Nous Research's assistant app, and I've been using it as the reference design for most of this. In Muse each side chat holds one goal. The agent writes it, checks in on a schedule, closes it with evidence, and you can see a dated log of what it did at each check-in. I wanted to steer the work without maintaining another task list.

I decided early to copy that lifecycle and nothing else. Grok Bot has per-item checklists and event triggers, OpenClaw has hooks, Dots has approval gates. I left those out because I didn't have a use for them yet. At one point my own agent suggested a fallback mode "in case the judge is unavailable" and I asked it whether Muse has a fallback. It doesn't. We don't either.

What Hermes already had

Hermes Agent, the open-source runtime we run on, has one "profile" per agent. A profile is a process, a model, a chat bot, a set of rules. We run eleven. Before last week, Hermes already had most of the goal plumbing:

A /goal command, where you type an objective and it's stored on the chat session. A judge: after every turn a small model (we use Gemini Flash) reads the session and returns done, progress, blocked, or unachievable. A turn budget, 20 by default, so a goal can't spin forever. A heartbeat, which is a recurring "anything to do?" tick inside a session. And cron jobs that can attach to an existing conversation instead of starting a blank one.

The one thing missing: only the human could create a goal. If an agent noticed "this needs following up on" it had nowhere to put that. The next turn it wouldn't remember.

What we added

A goal tool, turned on per profile with goals.agent_tool: true. Three actions:

goal(get)
goal(create, objective, acceptance_criteria, verification, max_turns?)
goal(complete, evidence)

The system prompt says: at the start of any turn that commits you to a multi-step outcome, call get. If nothing's set, create before you start working. Keep going until complete passes. If complete is refused, that means keep working, not rephrase the evidence.

Rules that came out of copying Muse and one round of critique:

The change is PR #134448 upstream, still a draft. The exact revision we run is pinned in the plugins repo README.