← Writing

Stop growing the system prompt. Coach the agent instead.

We moved behaviour rules out of the system prompt and into one-line, per-turn nudges picked by a cheap judge. Here is what we measured.

We run an eve agent in production at Featured. For a while, every time it misbehaved in some specific situation, the fix was another rule in the system prompt. Each rule fixed the one thing and quietly regressed something else, and the prompt kept growing.

Most of those rules only matter on some turns. Paying for them on every turn, and risking collisions between them on every turn, felt wrong. So yesterday I wrote up what we did instead and posted it to the eve discussions. This is the longer version.

The idea

A cheap judge model reads the incoming message, or the turn that just finished. When a rule applies, the agent gets one short line of guidance for that turn only. When no rule applies, the agent pays nothing.

We started calling it a coach. It does not rewrite answers and it does not block them. It says "do X this time" and gets out of the way.

10/10
Picked the right tool on turn 1
0/10 without the nudge
270 ms
Judge latency, median
one call per message
0%
Cache reuse with a system-scoped nudge
94–97% when ephemeral
10/10
Prompt rule broken in sampled answers
it was in the prompt all along

What we measured

1. Does a one-line nudge even reach the model?

Before measuring anything real, we checked that a single line appended to a turn is actually read, across multi-turn conversations. We used a nonsense rule, "always include the word banana", so there was no baseline behaviour to confound it.

PlacementTurns that complied
No nudge (control)0/8
Appended as user-role context for the turn8/8
System-scope instruction for the turn8/8
Client-supplied ephemeral context on send8/8

Every placement works. That left us free to choose placement on cost, which turned out to matter a lot.

2. Does it change a real behaviour?

On the first turn of a new conversation, our agent often reached for generic web search when the person wanted our own domain search. We sent a one-line nudge, only when a judge said the message was that kind of ask.

Picked the domain tool on turn 1, out of 10 conversations
  1. Run 1, no nudge0
  2. Run 1, nudged10
  3. Run 2, no nudge1
  4. Run 2, nudged10

Run 2 used a fuller search index, which is why the control improved slightly. The nudged runs were perfect both times.

3. What does it cost in prompt caching?

This decided the placement for us. Putting per-turn content into system scope dropped prefix reuse on the first model call of that turn from roughly 94 to 97 percent down to 0 percent. It recovered on later calls within the turn, but that first call is the one the person is waiting on.

Appending the nudge as ephemeral context at the end of the conversation kept reuse intact.

Placement of the nudgePrefix reuse on the turn's first call
None94–97%
System scope0%
Ephemeral context, end of transcript94–97%

4. Can a judge decide when to coach, cheaply and safely?

One batched classification call per message, made before the turn starts. We set a high threshold because a wrong nudge costs more than a missed one. A miss just falls back to default behaviour.

SetLabels correctFalse nudgesMissesLatency p50 / p90
Tuning (60 cases)57–58/6004–5270 ms / 470 ms
Fresh holdout, run once (22 cases)18/2214277 ms / 455 ms

Under 300 milliseconds at the median, and the errors lean the safe way.

5. Instructions alone don't always hold

One prompt rule, "offer choices rather than ending on a question in prose", was broken in 10 of 10 sampled answers. It was in the system prompt the whole time.

That is the kind of rule we are now moving to a deterministic check after the turn, with a nudge on the next turn only when it was broken.

What didn't work

Not everything works. One nudge we tested was clearly received and then ignored. I don't have a tidy explanation for it yet.

The lesson I took: each nudge has to earn its place with its own measurement. "We added a rule" means nothing until you have counted how often it fired and whether the next answer complied.

What first-class support could look like

If this pattern belongs in the framework rather than bolted on beside it, here is what I would want.

Open questions

I asked the eve team four things, and I would ask anyone running agents the same:

  1. Is this a framework primitive, or are instructions plus client context the intended way?
  2. Should pre-turn hooks be able to see the incoming message?
  3. Would you guarantee a cache-safe placement for per-turn context?
  4. Is anyone else running judge-in-the-loop guidance like this, and what did you measure?

If you have numbers, I want to see them. The discussion is the best place to reply.