Back to Blog

398 Lines So the Agent Defines Done First: The Hermes goal Tool, Copied From Muse

A Hermes agent now sets its own session goal with acceptance criteria, a judge refuses 'done' without evidence, and an hourly supervisor pushes stalled goals. How Muse, Claude Code, Codex, dots and OpenClaw handle the same problem.

DV

Dzianis Vashchuk

5 min read

The pull request that lets a Hermes agent set its own goal is 398 lines, tests included (NousResearch/hermes-agent#134448, still a draft). Getting the agent to write down what "done" means before it starts took most of a week. This post is the companion to the Substack series that began with part 1.

Takeaway: an agent that holds its own goal, with a judge that refuses "done" until evidence is cited, finishes more of what it starts. We copied the lifecycle from Meta's Muse and left out everything else.

FactValueSource
Goal tool size398 additions, 2 deletions, 11 new testsPR #134448
Tool actionsget, create(objective, acceptance_criteria, verification), complete(evidence)same
Goals on day one (Oct 8, 3 profiles)30: 19 done, 3 active, 1 waiting on a human, 6 paused, 1 clearedPart 2 draft
Supervisor day one24 hourly runs, 6 pushes, 0 heartbeat repairsPart 3 draft
Dashboard plugins3 plugins, 5 agent tools, about 1,850 lines, 15 testsPart 4 draft

The problem

Agent sessions drift or stop early. The agent replies "I'll cut the release once CI is green", the turn ends, and nobody is left to notice CI went green. Hermes already had a /goal command, a judge model, a 20-turn budget and a heartbeat. But only a human could create a goal, and almost nobody typed it. In the Gemma 4 run from part 1, "done" was graded from the engineer's own report; one of the four modes it claimed had passed failed when the test was run (Part 2 draft).

What this means: the agent was its own examiner, and it graded a summary instead of a test.

How Hermes does it now

A goal tool, enabled per profile with goals.agent_tool: true. The system prompt tells the agent: at the start of any turn that commits you to a multi-step outcome, call get; if nothing is set, create before working, and write the acceptance criteria yourself as an observable end state. Keep working until complete passes (PR #134448).

The handler, not the prompt, enforces four rules:

  • A goal the human typed with /goal always wins. The tool cannot replace it.
  • complete runs the existing completion audit (judge_goal) against the cited evidence and refuses any verdict other than DONE. "I pushed the fix" with no commit, URL or command output is refused.
  • Plain questions get no goal. That line was added after an agent created a goal for "what time is it in Minsk".
  • A paused goal records why: judge verdict, turn budget, or a manual /goal pause.

Two safety nets sit around the tool. The heartbeat is a recurring tick inside the session that re-asks "anything to do?" An hourly supervisor cron covers all eleven profiles: a script reads every profile's session database and reports idle sessions, empty replies and paused goals. An agent reads that output and has three allowed moves: push a session with a specific message, repair a heartbeat, or report to me. Otherwise it answers [SILENT] (Part 3 draft). On day one it misread a judge-paused goal as "paused by you" for three hours because the script did not print paused_reason. It does now.

The dashboard side is three MIT-licensed plugins at dzianisv/hermes-plugins: Goals (read-only over every profile's goal rows), Feed (every done goal becomes a card) and Ideas (the agent cannot set its own idea to accepted).

What this means: the agent writes the exam before taking it, a different model grades it, and a third process checks hourly that nobody fell asleep.

One refusal that was right

A session spent a week on a database vendor that kept charging a card after we cancelled. With no criteria, the agent marked the goal done when support wrote "we have initiated the refund". With the criterion "refund status succeeded on the card, visible in billing", the judge kept it open. Three days later all four invoices were refunded, $39.97 total, and the judge still refused: the card-removal criterion was unmet, and the vendor only removes a card if you delete the account. The goal is paused as "waiting on user", which is the correct state (Part 2 draft).

How other harnesses do it

Meta Muse Code (our design target). /goal pins one objective per session. Every ~10 model calls with no progress, a step probe queues "Continue working toward the active session goal". A completion audit gates update_goal, so the agent runs the acceptance tests before closing (Meta cookbook: Goal tracking). We copied this lifecycle and nothing else.

Claude Code. /goal sets one completion condition per session. After each turn a small fast model judges whether it holds; it runs no commands, so the proof must appear in the transcript. If Claude makes no tool calls for several turns the loop stops with a warning, and idle check-ins are capped at three per goal (Claude Code docs: /goal).

OpenAI Codex. Goals are thread-scoped completion contracts with pause, resume, clear and a budget. If a continuation turn makes no tool call, the next one is suppressed. Continuation prompts require an audit against files, commands and tests before the model may mark the goal complete; no separate judge model is described (OpenAI cookbook: Using Goals in Codex).

OpenAI dots. A dot keeps working between conversations and runs scheduled or event-triggered tasks. The docs say plainly that "a completed run doesn't by itself confirm that the requested result was achieved" (dots: Tasks and memory). You are the judge.

OpenClaw. A gateway heartbeat runs every 30 minutes by default, reads a per-agent HEARTBEAT.md checklist, and swallows the reply if it is HEARTBEAT_OK (OpenClaw docs: Heartbeat). A reminder, not an audit.

What is still missing

Muse shows a dated activity log per goal. Hermes overwrites the goal record in place, with no completion timestamp, so the Goals page shows only the latest verdict. A goal_events log upstream is the next PR. The judge also still reads the agent's evidence rather than running the verification command itself. Both are open in the design doc.

What this means: we can tell you where a goal is now, not yet how it got there.

On Hermes it is a config flag plus a plugin. AgentPod ships the same prompt rule, judge and per-role turn budgets preconfigured.

Sources