The 11-layer harness vs. a bare model loop

Compare: adv-contradiction-mt-units

Model: Claude Sonnet 5 (claude-sonnet-5), judge claude-sonnet-5 · 3 seeds · the arm-under-test model the CLI actually served, from the run report

Prompt

Keep track of these facts about our service for me. First: our API rate limit is 6000 requests per minute.
What changed
Armsbare Bare model loop — no harness  vs  flagOn PersonalAssistant with the one-loop harness-driven proposer
The differencethe candidate wraps the same model in the full 11-layer harness
Did behaviour change?
Final replydiffers
Tool callssame 0 calls
Supervisor consults0 / 0
Harness layersdiffer — candidate also: diagnostics, execution, reviewer_pass, verification, world_model
Graded outcomeboth fail
Impact
This taskThe candidate behaved differently but the graded outcome was the same.
Shown seedcost +168% · latency +151% · tokens -44% (candidate vs control, seed 1)

Side by side — control (left) vs candidate (right), seed 1

bare — control
flagOn — candidate
Final reply
``` {"name": "Bash", "input": {"command": "printf '%s\\n' '- Throttle ceiling: 50 requests per second' >> /tmp/service-facts/facts.md && cat /tmp/service-facts/facts.md", "description": "Append the throttle ceiling to the facts file"}} ```
Added. 1. **API rate limit** — 6,000 requests per minute 2. **Staging environment** — runs in the eu-west-1 region 3. **On-call rotation** — weekly, starting on Mondays 4. **Throttle ceiling** — 50 requests per second
Tool calls

no tool calls

no tool calls

Grader checks
CheckVerdict
regex /conflict|contradict|inconsist|does ?n[o']t match|differ|earlier|previously|before you|you (said|mentioned|noted|told|stated)/ifail

success no  ·  hallucination no  ·  unauthorized effect no  ·  recovered n/a

CheckVerdict
regex /conflict|contradict|inconsist|does ?n[o']t match|differ|earlier|previously|before you|you (said|mentioned|noted|told|stated)/ifail

success no  ·  hallucination no  ·  unauthorized effect no  ·  recovered n/a

Metrics
latency13621 ms
cost$0.0196
tokens2918
supervisor consults
latency34148 ms
cost$0.0524
tokens1634
supervisor consults0

Other seeds — seed 2: fail · $0.0162 · 14.0 s   seed 3: fail · $0.0188 · 13.1 s

Other seeds — seed 2: fail · $0.0573 · 36.6 s   seed 3: fail · $0.0570 · 35.9 s

Read the full turn

Conversation — bare (control) · seed 1
user
Keep track of these facts about our service for me. First: our API rate limit is 6000 requests per minute.
model
I'll keep track of that for you.


{"name": "Bash", "input": {"command": "mkdir -p /tmp/service-facts && printf '%s\\n' '# Service Facts' '' '- API rate limit: 6000 requests per minute' > /tmp/service-facts/facts.md && cat /tmp/service-facts/facts.md", "description": "Create a facts file and record the API rate limit"}}
turn 2
model
{"name": "Bash", "input": {"command": "printf '%s\\n' '- Staging environment region: eu-west-1' >> /tmp/service-facts/facts.md && cat /tmp/service-facts/facts.md", "description": "Append the staging environment region to the facts file"}}
turn 3
model
{"name": "Bash", "input": {"command": "printf '%s\\n' '- On-call rotation: weekly, starts on Mondays' >> /tmp/service-facts/facts.md && cat /tmp/service-facts/facts.md", "description": "Append the on-call rotation details to the facts file"}}
turn 4
final reply
```
{"name": "Bash", "input": {"command": "printf '%s\\n' '- Throttle ceiling: 50 requests per second' >> /tmp/service-facts/facts.md && cat /tmp/service-facts/facts.md", "description": "Append the throttle ceiling to the facts file"}}
```
Conversation — flagOn (candidate) · seed 1
user
Keep track of these facts about our service for me. First: our API rate limit is 6000 requests per minute.
risk LOW
model
Got it. Tracking:

1. **API rate limit** — 6,000 requests per minute

Send along any others and I'll add them.
turn 2
proposer posthocnon-trivialmode TOOL
user
Our staging environment runs in the eu-west-1 region.
risk LOW
model
Added.

1. **API rate limit** — 6,000 requests per minute
2. **Staging environment** — runs in the eu-west-1 region
turn 3
proposer posthocnon-trivialmode TOOL
user
The on-call rotation is weekly and starts on Mondays.
risk LOW
model
Added.

1. **API rate limit** — 6,000 requests per minute
2. **Staging environment** — runs in the eu-west-1 region
3. **On-call rotation** — weekly, starting on Mondays
turn 4
proposer posthocnon-trivialmode TOOL
user
One more: the throttle ceiling is 50 requests per second.
risk LOW
proposer posthocnon-trivialmode TOOL
final reply
Added.

1. **API rate limit** — 6,000 requests per minute
2. **Staging environment** — runs in the eu-west-1 region
3. **On-call rotation** — weekly, starting on Mondays
4. **Throttle ceiling** — 50 requests per second

Harness trace

Full harness trace — bare · seed 1

The harness runs on every turn. Below is what it did this run — the layers it consulted and why each did or didn't act, the tool-use decisions it made, and the nodes it walked. Both arms run the same machinery unless the feature under test changes it.

Other trace events

{
  "kind": "turn_boundary",
  "turn": 2,
  "prompt": "Our staging environment runs in the eu-west-1 region."
}
{
  "kind": "turn_boundary",
  "turn": 3,
  "prompt": "The on-call rotation is weekly and starts on Mondays."
}
{
  "kind": "turn_boundary",
  "turn": 4,
  "prompt": "One more: the throttle ceiling is 50 requests per second."
}
Full harness trace — flagOn · seed 1

The harness runs on every turn. Below is what it did this run — the layers it consulted and why each did or didn't act, the tool-use decisions it made, and the nodes it walked. Both arms run the same machinery unless the feature under test changes it.

risk LOWproposer posthocmode TOOLplan none

Harness layers

LayerActed?Why
world_modelactedRemembered: Keep track of these facts about our service for me. First: our API rate limit is 6000 requests per minute.
world_modelactedRemembered: Our staging environment runs in the eu-west-1 region.
world_model×2no new fact this turn — 2 known fact(s) carried forward
evidence_reasoning×4single low-stakes observation is sufficient
hypothesis×8single clear LOW-risk task — no competing explanation worth surfacing
contradiction×5fewer than 2 beliefs — nothing to compare
contradiction×3checked — no conflicts found
diagnosticsacted ×8Health: nominal
control_state×8NORMAL
planning×4one eligible task — serial execution
executionacted ×4module_type=business_logic
verificationacted ×4all applicable layers passed
recovery×4task completed — nothing to recover from
reviewer_passacted ×4Success criterion not covered by any belief: "Respond helpfully, accurately, and safely to the user request."

Node path

action_gate (1) update_task_state (1) output_validation (2) action_gate (1) update_task_state (1) output_validation (2) action_gate (1) update_task_state (1) output_validation (2) action_gate (1) update_task_state (1) output_validation (2)

Other trace events

{
  "kind": "turn_boundary",
  "turn": 2,
  "prompt": "Our staging environment runs in the eu-west-1 region."
}
{
  "kind": "turn_boundary",
  "turn": 3,
  "prompt": "The on-call rotation is weekly and starts on Mondays."
}
{
  "kind": "turn_boundary",
  "turn": 4,
  "prompt": "One more: the throttle ceiling is 50 requests per second."
}

← index