Guardrails for coding agents · Part 3
Your test suite is slowing the agent down
Moving tests, type checks and evals out of the agent's way, the throughput it bought, and the limits of parallel agents on one machine.
The most expensive guardrails are not the ones that block. They are the ones that wait.
A human developer runs the test suite a few times a day and reads email while it runs. A coding agent commits every few minutes as a checkpoint, and a suite in the commit path turns every checkpoint into a five-minute pause. Multiply by a session of fifty commits and you have paid for a suite run per commit and shipped nothing extra for it. Part 2 set a twenty-second budget for the pre-commit hook. This part is about where the rest of the verification went.
Run the suite after the push
In mid September the rule became: verify after pushing, never before committing. Commit and push first. A hook on the push then runs type checks, the full test suite, the agent compile check, lint and the knowledge-graph link check in the background. It stays silent when everything passes and wakes the agent with the failures when something does not. The fix lands in the next commit.
Two details made this work rather than just move the wait:
- A newer push supersedes the suite in flight. Only the branch tip matters, the same way a CI concurrency group works. Without this, three suites ran at once and exhausted the local database's connections, failing each other.
- The agent is woken, not polled. The hook exits non-zero with the failures on standard error only when something fails. The agent never spends a turn asking whether the suite is done.
Continuous integration still runs the same suite on the pushed commit. The hook is for the agent's loop, not a replacement for the shared gate.
Build first, test last
The same month the instruction for feature work became blunter: write the code and commit as you go; run tests, type checks, builds and local servers once, near the end, when the work is close to ready. Applies to subagents too, and they are told explicitly.
The reason was observed, not theoretical. Mid-build verification runs were the single largest line item in a session's wall-clock time, and most of what they found was in code that was about to change again anyway. A few weeks later the rule got its corollary: a red test mid-stream is a later fix, not a blocker, and nobody apologises for it. Save the full verification pass for just before the pull request.
And then: do not watch CI. After opening or pushing a pull request, link it and stop. No polling, no "waiting for the type check before merging". The agent's habit of sitting on a green light was costing more than the occasional red one.
Which tests earn their place
Verification moved later, and it also got smaller, because a lot of it was not worth running at any time.
Tautological tests were the first cut. A test that lists every tool in a registry, or asserts the exact copy string from a messages file, or snapshots a config object, fails on every legitimate change and catches nothing. When adding a registry entry forces an edit to a test that only lists entries, the test is tautological: turn it into a property check ("every tool that can spend money is denied to anonymous callers") or delete it. This one went into the global instructions for every repository.
Compile-time checks over tests where the type system can prove the thing. An exhaustive table bound with satisfies, so that adding a union member is a compile error until every surface handles it, replaces a test that would have checked the same coverage more slowly and less reliably.
Evals, not yet. For a product whose shape is still changing weekly, we stopped writing and extending model evaluation suites. The owner's words: we are too early for evals, we do not even know what the product is. A handful of targeted prompts and a cheap judge over a few cases gives the signal. The suite comes back when the product settles.
The limits of parallel agents
The throughput story has a physical ceiling, and we found it by hitting it.
On one build, five agents ran in parallel, each in its own git worktree, each with its own installs, type checks, test suites, dev servers and browsers. The orchestrating session's shell was killed by the operating system for memory, exit code 137, in the middle of a merge. The machine has 39 GB.
The rules that came out of that afternoon are not clever, which is the point:
- At most three heavy agents at once.
- Never two full test suites at the same time. Cap the test runner's workers.
- Stop every dev server, browser and watcher the moment its work is done.
- Check free memory before starting heavy work.
- Save in-progress work as commits, so a kill loses nothing.
Two days later a second limit showed up, this time in wall-clock time rather than memory. The owner's note, verbatim because it is the right level of bluntness: subagents are slow, my man. A subagent starts cold. It re-reads context, briefs and documents that the main session already has. For a small fix, a copy change, a single screenshot, that cold start costs more than the work. So: small and medium tasks happen in the main session. Subagents are for genuinely parallel heavy work and long runs. And when a background agent already has partial output on disk, show it immediately rather than waiting for its report.
Never block the turn on a deploy
One more habit had to go. The agent would push, then sit in the foreground waiting for the hosting platform's build to finish before doing anything else. The owner caught it: you were stuck.
The rule now is to check a deployment once, report its state, and put any wait in the background. And the inverse rule, which turned out to matter more: if you are waiting for a deploy anyway, prepare the next batch of changes in a second worktree while it runs, and fold them into the next deploy. Deploy waits became a scheduling opportunity rather than dead time.
What carried over
- Verification after the push, in the background, waking the agent only on failure.
- Build first, verify once, close to the end. A red test mid-stream is a later fix.
- Cut tests that restate the code. Property checks and compile-time proofs instead.
- Three heavy agents, one suite at a time, servers stopped when done. The ceiling is the machine.
- Small work in the main session. Subagents for parallel heavy work only.
- Never wait in the foreground. Check once, report, prepare the next batch.
Everything in this part is about speed. The next part is about the week speed turned into damage, and the guardrails on the human side of the loop that came out of it.