Every agent that survives a multi-hour task is wrapped in something. A task list it keeps returning to. A rule that says don’t stop because a milestone felt like a good place to report. A verifier that refuses to accept “done” without evidence. We call this the harness, and for hard tasks it often matters more than the model inside it.11“Harness” here means everything around the model call: the loop, the tools, the prompts, the files the agent reads and writes.
That raises a question I keep coming back to: how much of the harness is knowledge the model could hold itself?
01#Three kinds of scaffolding
Looking at the harnesses I’ve written over the past year, the pieces fall into roughly three groups, nested around the model like the layers of an onion.
- Memory: checklists, scratch files, summaries that survive a context reset. These compensate for a finite window.
- Discipline: rules about when to stop, when to verify, when to ask. These compensate for habits the model learned elsewhere.
- Judgment: routing a subtask to the right tool, deciding that a result is good enough. These encode taste.
A harness rule is a correction written down once and paid for on every call.
Memory seems like infrastructure; it will probably always live outside the weights. Discipline is the interesting middle. Judgment is where it gets hard.
02#Pricing a rule
One way to see why internalising matters: every rule in the prompt costs tokens on every call, and attention spent reading the rule is attention not spent on the task. If a harness has rules of average length , and an episode makes calls, the overhead is roughly
For a long run ( in the hundreds) that adds up fast. A rule the model has internalised costs nothing at inference time.
03#What removal might look like
The experiment I want to run: start with a harness that works, delete one rule at a time, and measure how often a long task still finishes correctly, first for the base model and then for a model trained on trajectories from the full harness.
Task success as harness rules are removed (illustrative)
Success rate on long tasks, by number of rules removed
- Base model
- Trained on harness traces
| Base model | Trained on harness traces | |
|---|---|---|
| 0 | 78% | 80% |
| 1 | 71% | 79% |
| 2 | 63% | 76% |
| 3 | 52% | 73% |
| 4 | 44% | 69% |
| 5 | 35% | 62% |
| 6 | 29% | 55% |
If the trained curve stays flat while the base curve falls, the rules on the flat stretch have moved into the weights.
04#A tiny example
The simplest discipline rule I use looks like this:
def should_continue(turn, tasks, attempts): # A text-only end of turn is a report, not proof of completion. if turn.stop_reason == "end_turn" and tasks.open(): return attempts < 3 return FalseIt works. But it is also a sign that the model, left alone, believes it is finished before it is. If a model internalised this, the rule would become dead code.
05#Which rules move first
Share of runs where removing the rule hurt (illustrative)
Runs that got worse without the rule
- Base model
- Trained
| Base model | Trained | |
|---|---|---|
| Task list | 64% | 61% |
| Stop rule | 48% | 12% |
| Verify rule | 41% | 15% |
| Tool routing | 22% | 18% |
| Scaffolding | Lives in | Could move to weights? |
|---|---|---|
| Task list | files | unlikely |
| Stop / continue rules | prompt + loop | plausibly |
| Verification habits | prompt | plausibly |
| Tool routing | code | partly |
This is a placeholder post written to preview the blog layout. Real writing will replace it.