Chapter 9 of 15

Done Means Green

Freddy asked an agent to add the search box. It replied that it was done, and the build was red. He has had that reply before. This chapter is how he stops taking it on trust: a goal that says what finished means, a check that proves it, and a limit on what it may cost to try.

The problem

“Done” is the agent's claim, and agents are optimistic. They stop when they believe the task is complete, which is not when it is. The other failure is the opposite: a big task that needs ten steps, where you sit and say “continue” ten times.

The hard way

Read the code, run the tests yourself, paste the failures back, repeat. It works, and you are the check command. For the long task, you are also the loop. And nothing in between stops an agent that has been given a hard problem from spending an afternoon of your budget on it.

Goals

A goal keeps the agent working, turn after turn, until a larger objective is reached. For example: All tests pass and the checkout flow handles discounts.

  1. Write the goal in the Context tab and press Start goal. The agent gets the goal as its next message, with instructions to take one concrete step at a time and verify it.
  2. After each turn, Workspace sends it on automatically, until one of three things happens:
    • the agent ends a reply with GOAL ACHIEVED: the goal is Achieved;
    • the agent ends a reply with GOAL BLOCKED because it needs you, or a turn fails: the goal is Paused;
    • the turn budget runs out (10 automatic turns by default, set in Settings → Agents & MCP): the goal stops.
    You get a notification in each case.
  3. Pause stops after the current turn. Resume continues with a fresh budget. Clear removes the goal.

Messages you queue yourself are sent before the goal continues. And while a goal is active, permission requests still stop and wait for you, so choose the permission mode with that in mind. A goal in Ask for approval mode is a goal that will pause at every edit.

Budget

A spending limit for the thread, in US dollars. The Budget section shows what the thread has cost so far, with a bar when a limit is set. Type an amount and press Set (or Enter); Remove takes the limit away.

  • At 80% of the limit, Workspace notes it in the conversation.
  • At the limit, a running goal pauses and asks you, and a goal cannot be started or resumed until you raise the limit. Messages you send yourself still go.
A limit that does nothing for some agents. Costs are the ones the agent reports with each turn. Claude Code reports them; agents that do not report costs count as free, so a limit has no effect on them. If you rely on a budget, check that the agent in the thread reports a cost at all.

Checks after each turn

Done means green. Give the project a check command: tests, lint, or both, such as cargo test or vendor/bin/pint --test && php artisan test. Workspace suggests one from the project's files; Use … fills it in, and Run now tries it. Freddy's is npm test.

  • After every turn that changed files (whether a turn changed files is judged against the snapshot taken before your message reached the agent, so a quick agent cannot slip a change past the checks), the command runs in the thread's folder, with your login shell and CI=1. A turn that changed nothing is not checked.
  • The thread stays busy while the checks run. Checks passed or Checks failed appears in the conversation; click it for the output.
  • When they fail, the end of the output goes back to the agent to fix (Sent the failures to the agent), and the checks run again. After Settings → Agents & MCP → Automatic fixes when checks fail (2 by default) the thread stops, marked failed, and you are told. 0 only reports the failure. No automatic fix is sent when you have queued a message or the budget is used up.
  • A stronger fix. When a first automatic fix did not make them pass, the next one gets the agent's Escalation model (Settings → Providers, for example opus) and its highest effort; the thread goes back to its own model afterwards. With no escalation model set, only the effort goes up.
  • Stop stops the checks too. A check is stopped after 20 minutes.
  • Notifications, Félagi reports and wait_for_thread wait for the checks, so finished means the checks passed.

The command applies to every thread in the project. Leave it empty to turn checks off.

Read the escalation setting as a way to spend money where it helps: a cheaper model for the easy turns, and the strong one only when the first fix did not work. And notice what the last bullet gives you in practice. Freddy can start a goal, leave, and the notification that says a thread finished means the tests passed, not that the agent stopped typing.

A check is only as good as its command. If npm test covers nothing in the code the agent changed, green means nothing. Checks do not replace review (chapter 5); they make sure the review starts at something that runs.
Try it: set a check command for a project and Run now. Then give a thread a small task that you know will break a test, and watch Checks failed go back to the agent. Set a budget of a few dollars while you are there.

What you learned

  • How a goal keeps an agent going, and the three ways it ends
  • That permission requests still wait for you while a goal runs
  • How a budget works, and that it has no effect on agents that do not report costs
  • How the check command runs after each turn, and what happens when it fails
  • How the stronger second fix uses the escalation model
  • Why a finished notification means the checks passed
Next: in Chapter 10 you give the agent a browser: what it can see, what it can click, and how a flow it tried becomes a journey that is replayed like a test.