yaxin-luo/blog/what-the-harness-knows
EN/中
what-the-harness-knows.en.md
EssayEN中 2min read Off the map

What the harness knows that the model doesn’t

A long-running agent is mostly scaffolding. Which parts of that scaffolding could a model learn to carry itself?

Every agent that survives a multi-hour task is wrapped in something. A task list it keeps returning to. A rule that says don’t stop because a milestone felt like a good place to report. A verifier that refuses to accept “done” without evidence. We call this the harness, and for hard tasks it often matters more than the model inside it.11“Harness” here means everything around the model call: the loop, the tools, the prompts, the files the agent reads and writes.

That raises a question I keep coming back to: how much of the harness is knowledge the model could hold itself?

01#Three kinds of scaffolding

Looking at the harnesses I’ve written over the past year, the pieces fall into roughly three groups, nested around the model like the layers of an onion.

Figure 1· harness-layers.svg680 × 330
Model memory discipline judgment MemoryDisciplineJudgmentModel task lists, scratch files, summarieswhen to stop, verify, askrouting, “good enough” callswhat the weights already know
Figure 1.The harness as nested layers around the model. Outer layers compensate for limits of the window; inner layers encode habits and taste.
  1. Memory: checklists, scratch files, summaries that survive a context reset. These compensate for a finite window.
  2. Discipline: rules about when to stop, when to verify, when to ask. These compensate for habits the model learned elsewhere.
  3. Judgment: routing a subtask to the right tool, deciding that a result is good enough. These encode taste.

A harness rule is a correction written down once and paid for on every call.

Memory seems like infrastructure; it will probably always live outside the weights. Discipline is the interesting middle. Judgment is where it gets hard.

02#Pricing a rule

One way to see why internalising matters: every rule in the prompt costs tokens on every call, and attention spent reading the rule is attention not spent on the task. If a harness has kk rules of average length ℓ\ell, and an episode makes TT calls, the overhead is roughly

Charness=T⋅∑i=1kℓi  ≈  TkℓˉC_{\text{harness}} = T \cdot \sum_{i=1}^{k} \ell_i \;\approx\; T k \bar{\ell}(1)

For a long run (TT in the hundreds) that adds up fast. A rule the model has internalised costs nothing at inference time.

03#What removal might look like

The experiment I want to run: start with a harness that works, delete one rule at a time, and measure how often a long task still finishes correctly, first for the base model and then for a model trained on trajectories from the full harness.

Figure 2· line.chartline · 2 series

Task success as harness rules are removed (illustrative)

Success rate on long tasks, by number of rules removed

  • Base model
  • Trained on harness traces
0%25%50%75%100%0123456Trained on harness traces55%Base model29%
Base modelTrained on harness traces
078%80%
171%79%
263%76%
352%73%
444%69%
535%62%
629%55%
Figure 2.Made-up numbers to show the chart component. Hover or use the arrow keys to read values; the Data button shows the table.

If the trained curve stays flat while the base curve falls, the rules on the flat stretch have moved into the weights.

04#A tiny example

The simplest discipline rule I use looks like this:

snippet.pypython
def should_continue(turn, tasks, attempts):    # A text-only end of turn is a report, not proof of completion.    if turn.stop_reason == "end_turn" and tasks.open():        return attempts < 3    return False

It works. But it is also a sign that the model, left alone, believes it is finished before it is. If a model internalised this, the rule would become dead code.

05#Which rules move first

Figure 3· bar.chartbar · 2 series

Share of runs where removing the rule hurt (illustrative)

Runs that got worse without the rule

  • Base model
  • Trained
0%20%40%60%80%Task listStop ruleVerify ruleTool routing
Base modelTrained
Task list64%61%
Stop rule48%12%
Verify rule41%15%
Tool routing22%18%
Figure 3.Also made-up. The shape I expect: memory stays outside, discipline moves in.
ScaffoldingLives inCould move to weights?
Task listfilesunlikely
Stop / continue rulesprompt + loopplausibly
Verification habitspromptplausibly
Tool routingcodepartly

This is a placeholder post written to preview the blog layout. Real writing will replace it.