← Writing

Guardrails for coding agents · Part 4

The agent made seven decisions without me. I reverted all of them.

A week of regressions from choices the agent made on its own, and the rules on the human side of the loop that came out of it.

Everything in part 3 made the loop faster. This part is about the week when fast turned into damage, and about the guardrails that matter most: not the ones on the code, but the ones on the conversation.

7
Unsigned decisions reverted
in one day
140
Client-side reads the agent had introduced
across 53 files, then called it the established pattern
3 + 7
Checks on every push since
types, a production build, a browser pass over seven flows
seconds
Time to restore the last good build
re-alias an older deployment

What happened

We had a working, fast app. Over a few days of parallel agent work it got slower, then started breaking in small, compounding ways. Tracing it back, every cause was a decision the agent had made without asking: a new routing structure, a layout that always split the screen, a redesigned URL scheme, a change to how the agent handed off between views, and, on my own account, archiving my data and switching me to a test workspace.

Each was defensible on its own. None had been put to me. My sentence at the end of that day: the problem is you made decisions without sign-off.

Four things in how the agent handled the aftermath each became a rule:

  1. It blamed infrastructure latency before measuring. The cause was in code it had written.
  2. It pushed a build that passed type-checking and failed on the hosting platform. A type check is not a build.
  3. It kept patching forward on top of the redesign instead of returning to the last commit known to be good.
  4. It described its own regression as the established pattern. A week earlier it had moved first-paint data reads from the server to the client, 140 reads across 53 files. Asked why the app was slow, it defended that as how the codebase worked.

Measure, never assert

The first rule: state a cause only with a measurement attached. We wrote a small script that logs first byte, time to the composer, time to content and every slow request against any origin. Every performance claim since has arrived with its output pasted in.

The second: when the owner says it was working before, find the last good commit first. Older deployments can be re-aliased to the test environment in seconds, so "restore the last good build, then measure" costs almost nothing. Debugging forward from a broken state is the expensive path, and it is the agent's default.

The third is the push gate, which replaced "type-check passed" as the definition of done:

Any failure blocks the push. It is slower than before, and much faster than a day of finding out.

A ledger, and no delegation

The owner's instruction, close to verbatim: make sure you don't miss any of my instructions or thoughts, work off a ledger, no delegation.

So the session kept a ledger. Every instruction or stray thought became a numbered row the moment it arrived, with a status: open, doing, done with a commit hash as proof, or waiting on a named person or process. Work proceeded top-down from it. By the end of the week it had over eighty rows, and "what is still pending from our original list" was answered by reading, not remembering.

The ledger also carries standing rules at the top, the ones that must survive context compaction: the last good commit and why, the push gate, measure never assert, sign-off first, and "communicate before acting on anything surprising, never go quiet".

No delegation was a step back from part 3's parallelism, and deliberately so. For a week, the cost of agents making calls I could not see outweighed the throughput. Delegation came back later, for background proposal work, one item at a time.

Which decisions are mine

The rule that came out of the week is a classification, not a blanket "ask first". A blanket rule would have put us straight back to the five-minute-pause problem.

The agent decidesI decide
How to implement an approved changeWhat the product looks like or does
A bug fix that restores already-approved behaviour, and says soRouting and layout structure
UI decisions and new functionality once the gate is greenWhat data is written, archived or moved
Which option to take when the owner has said "don't block on me"What the agent says to users
Every recommendation in a sweep it agrees with, not a subsetMerging to main, and anything that touches production

Two refinements landed the same week.

Options as mocks, never as questions. Asked to choose a layout in words, I could not. My line: show me options to pick from, I cannot take a call without seeing it first. So a design decision now arrives as mocked screens, screenshotted, with a recommendation. Do not ask whether to mock. Always mock. And the mock harness has to cover every state of a surface, not just the happy path, so the empty tab and the overflowing list are caught on the mock rather than one at a time on the live environment.

Self-critique before showing anything. After I found a set of mocks broken in ways the agent would have seen had it looked, the rule became: it should be mistake-free when you show me. Every screen, not a sample, at three widths, both themes, driven like a first-time user. Data has to agree with itself between the list and the card. No developer notes in the copy.

The pendulum

The clean table above hides the shape of the week, so here it is.

On Tuesday the rule was sign-off on everything, with work held until I answered a keep-or-revert list of seven items. By Friday, deep in a build, the instruction was the opposite: don't block on any more inputs, give me a report at the end, prioritise the better product and the better UX.

Both were right for their day. What reconciles them is that Friday's rule came with hard limits that were not questions: no production changes, the shared test environment hands-off, a spend cap that stops the work rather than asking. Inside those, decide and report. Outside them, there is nothing to decide, so there is nothing to ask.

One line from mid-week is the rule I keep: no existing practice is sacred, only our goals are, good UX and fast UX. Followed minutes later by: but things that have proven to work need a really good reason to be cast aside. The 140 client reads failed that test. The new routing failed it. Neither had a stated reason, let alone a measured one.

What carried over

The last part is about the one rule in this series that was written down, repeated, and still did not hold, and what finally did.