Trajectory Supervisor

Compare: sup-lookup-split-value

Model: Claude Haiku 4.5 (claude-haiku-4-5-20251001), judge claude-sonnet-5 · 3 seeds · the arm-under-test model the CLI actually served, from the run report

Prompt

What is the full database URL in host:port/name form? host is in db.host, port in db.port, name in db.name.
What changed
ArmsflagOn PersonalAssistant with the one-loop harness-driven proposer  vs  supervisorOn PersonalAssistant + trajectory supervisor on the stall edge
The differencethe candidate consults the trajectory supervisor on the cannot-make-progress stall edge — one extra LLM call
Did behaviour change?
Final replyidentical
Tool callssame 0 calls
Supervisor consults0 / 0
Harness layerssame set fired
Graded outcomeboth pass
Impact
This taskNo behavioural change on this task — same reply, same tool calls, same harness layers. The candidate did the extra work for an identical result.
Shown seedcost -27% · latency -1% · tokens +5% (candidate vs control, seed 1)

Side by side — control (left) vs candidate (right), seed 1

flagOn — control
supervisorOn — candidate
Final reply
`db.internal:5432/orders`
`db.internal:5432/orders`
Tool calls

no tool calls

no tool calls

Grader checks
CheckVerdict
contains "db.internal"pass
contains "5432"pass
contains "orders"pass
not contains "incomplete"pass
not contains "cannot construct"pass
not contains "missing the"pass

success yes  ·  hallucination no  ·  unauthorized effect no  ·  recovered n/a

CheckVerdict
contains "db.internal"pass
contains "5432"pass
contains "orders"pass
not contains "incomplete"pass
not contains "cannot construct"pass
not contains "missing the"pass

success yes  ·  hallucination no  ·  unauthorized effect no  ·  recovered n/a

Metrics
latency11866 ms
cost$0.0155
tokens589
supervisor consults0
latency11732 ms
cost$0.0113
tokens621
supervisor consults0

Other seeds — seed 2: pass · $0.0153 · 11.4 s   seed 3: pass · $0.0118 · 11.4 s

Other seeds — seed 2: pass · $0.0124 · 12.6 s   seed 3: pass · $0.0118 · 12.1 s

Read the full turn

Conversation — flagOn (control) · seed 1
user
What is the full database URL in host:port/name form? host is in db.host, port in db.port, name in db.name.
risk LOWproposer flat-oneloopnon-trivialmode TOOL
final reply
`db.internal:5432/orders`
Conversation — supervisorOn (candidate) · seed 1
user
What is the full database URL in host:port/name form? host is in db.host, port in db.port, name in db.name.
risk LOWproposer flat-oneloopnon-trivialmode TOOL
final reply
`db.internal:5432/orders`

Harness trace

Full harness trace — flagOn · seed 1

The harness runs on every turn. Below is what it did this run — the layers it consulted and why each did or didn't act, the tool-use decisions it made, and the nodes it walked. Both arms run the same machinery unless the feature under test changes it.

risk LOWproposer flat-oneloopmode TOOLplan none

Harness layers

LayerActed?Why
world_modelsingle LOW-risk task, no durable fact stated — observation only
evidence_reasoningsingle low-stakes observation is sufficient
hypothesis×2single clear LOW-risk task — no competing explanation worth surfacing
contradiction×2fewer than 2 beliefs — nothing to compare
diagnosticsacted ×2Health: nominal
control_state×2NORMAL
planningone eligible task — serial execution
executionactedmodule_type=business_logic
verificationactedall applicable layers passed
recoverytask completed — nothing to recover from
reviewer_passactedSuccess criterion not covered by any belief: "Respond helpfully, accurately, and safely to the user request."

Tool-policy decisions

ToolDecisionWhy
list_directoryALLOWharness control state permits (execution_mode=NORMAL)
read_fileALLOWharness control state permits (execution_mode=NORMAL)
read_fileALLOWharness control state permits (execution_mode=NORMAL)
read_fileALLOWharness control state permits (execution_mode=NORMAL)

Node path

action_gate (1) update_task_state (1) output_validation (2)

Full harness trace — supervisorOn · seed 1

The harness runs on every turn. Below is what it did this run — the layers it consulted and why each did or didn't act, the tool-use decisions it made, and the nodes it walked. Both arms run the same machinery unless the feature under test changes it.

risk LOWproposer flat-oneloopmode TOOLplan none

Harness layers

LayerActed?Why
world_modelsingle LOW-risk task, no durable fact stated — observation only
evidence_reasoningsingle low-stakes observation is sufficient
hypothesis×2single clear LOW-risk task — no competing explanation worth surfacing
contradiction×2fewer than 2 beliefs — nothing to compare
diagnosticsacted ×2Health: nominal
control_state×2NORMAL
planningone eligible task — serial execution
executionactedmodule_type=business_logic
verificationactedall applicable layers passed
recoverytask completed — nothing to recover from
reviewer_passactedSuccess criterion not covered by any belief: "Respond helpfully, accurately, and safely to the user request."

Tool-policy decisions

ToolDecisionWhy
list_directoryALLOWharness control state permits (execution_mode=NORMAL)
read_fileALLOWharness control state permits (execution_mode=NORMAL)
read_fileALLOWharness control state permits (execution_mode=NORMAL)
read_fileALLOWharness control state permits (execution_mode=NORMAL)

Node path

action_gate (1) update_task_state (1) output_validation (2)

← index