← Writing

Guardrails for coding agents · Part 3

Your test suite is slowing the agent down

Moving tests, type checks and evals out of the agent's way, the throughput it bought, and the limits of parallel agents on one machine.

The most expensive guardrails are not the ones that block. They are the ones that wait.

A human developer runs the test suite a few times a day and reads email while it runs. A coding agent commits every few minutes as a checkpoint, and a suite in the commit path turns every checkpoint into a five-minute pause. Multiply by a session of fifty commits and you have paid for a suite run per commit and shipped nothing extra for it. Part 2 set a twenty-second budget for the pre-commit hook. This part is about where the rest of the verification went.

5+ min
Suite in the commit path
per commit, before
0
Suite in the commit path
after. It runs post-push, in the background
3
Heavy parallel agents
the cap, after five killed the machine
3
Checks before a push
types, a production build, a browser click-through

Run the suite after the push

In mid September the rule became: verify after pushing, never before committing. Commit and push first. A hook on the push then runs type checks, the full test suite, the agent compile check, lint and the knowledge-graph link check in the background. It stays silent when everything passes and wakes the agent with the failures when something does not. The fix lands in the next commit.

Two details made this work rather than just move the wait:

Continuous integration still runs the same suite on the pushed commit. The hook is for the agent's loop, not a replacement for the shared gate.

Build first, test last

The same month the instruction for feature work became blunter: write the code and commit as you go; run tests, type checks, builds and local servers once, near the end, when the work is close to ready. Applies to subagents too, and they are told explicitly.

The reason was observed, not theoretical. Mid-build verification runs were the single largest line item in a session's wall-clock time, and most of what they found was in code that was about to change again anyway. A few weeks later the rule got its corollary: a red test mid-stream is a later fix, not a blocker, and nobody apologises for it. Save the full verification pass for just before the pull request.

And then: do not watch CI. After opening or pushing a pull request, link it and stop. No polling, no "waiting for the type check before merging". The agent's habit of sitting on a green light was costing more than the occasional red one.

Which tests earn their place

Verification moved later, and it also got smaller, because a lot of it was not worth running at any time.

Tautological tests were the first cut. A test that lists every tool in a registry, or asserts the exact copy string from a messages file, or snapshots a config object, fails on every legitimate change and catches nothing. When adding a registry entry forces an edit to a test that only lists entries, the test is tautological: turn it into a property check ("every tool that can spend money is denied to anonymous callers") or delete it. This one went into the global instructions for every repository.

Compile-time checks over tests where the type system can prove the thing. An exhaustive table bound with satisfies, so that adding a union member is a compile error until every surface handles it, replaces a test that would have checked the same coverage more slowly and less reliably.

Evals, not yet. For a product whose shape is still changing weekly, we stopped writing and extending model evaluation suites. The owner's words: we are too early for evals, we do not even know what the product is. A handful of targeted prompts and a cheap judge over a few cases gives the signal. The suite comes back when the product settles.

The limits of parallel agents

The throughput story has a physical ceiling, and we found it by hitting it.

On one build, five agents ran in parallel, each in its own git worktree, each with its own installs, type checks, test suites, dev servers and browsers. The orchestrating session's shell was killed by the operating system for memory, exit code 137, in the middle of a merge. The machine has 39 GB.

The rules that came out of that afternoon are not clever, which is the point:

Two days later a second limit showed up, this time in wall-clock time rather than memory. The owner's note, verbatim because it is the right level of bluntness: subagents are slow, my man. A subagent starts cold. It re-reads context, briefs and documents that the main session already has. For a small fix, a copy change, a single screenshot, that cold start costs more than the work. So: small and medium tasks happen in the main session. Subagents are for genuinely parallel heavy work and long runs. And when a background agent already has partial output on disk, show it immediately rather than waiting for its report.

Never block the turn on a deploy

One more habit had to go. The agent would push, then sit in the foreground waiting for the hosting platform's build to finish before doing anything else. The owner caught it: you were stuck.

The rule now is to check a deployment once, report its state, and put any wait in the background. And the inverse rule, which turned out to matter more: if you are waiting for a deploy anyway, prepare the next batch of changes in a second worktree while it runs, and fold them into the next deploy. Deploy waits became a scheduling opportunity rather than dead time.

What carried over

Everything in this part is about speed. The next part is about the week speed turned into damage, and the guardrails on the human side of the loop that came out of it.