Second half of the strip-off isolation batch (SubagentsPi, loop_guard=on
subagent_nudges=on STRIP_THINKING=off KEEP_TOOL_RESULTS=0). Stopped by the
user at ~13:53 during trial 2 (headless-terminal, ~5 min in, no result —
its partial dir has no result.json; leaked env container was removed and
the mid-flight transcript strip was run manually). The fast-set half of the
batch had already completed and settled the A/B — see
runs/fast__qwen3.6-35b-a3b__20260708-115953/NOTES.md.
reshard-c4-data: PASS, reward 1.0, but 32m23s session vs 5m43s agent on the 07-07 baseline pass — this trial is the wall-clock exhibit for the SUBAGENT_NUDGES ceremony cost (new ledger item): 07-07 did scout+planner then self-served (335s session, 42 turns); this one ran the full nudged workflow — scout, planner ×2 + resume, worker, reviewer ×3 (the model re-called reviewer twice on its own after a clean first review) — 70 turns, 8 child calls ≈ 500s of child wall time, each child a fresh prefill on the shared llama-swap slot, plus ~450s of bash dead-ends (a fully-burned self-set 300s timeout + a FILE-NOT-FOUND detour) on top of the legitimate ~10k-file compress/verify work. Nudge dedup itself worked (exactly 3 nudges, one per stage). Same verdict either way — the ceremony bought nothing here but 5.8× the wall clock.
llama-local/qwen3.6-35b-a3bagentharnesses.subagents_pi:SubagentsPithinkingonreasoning budgetdirect (server/none)* — see journal for the authoritative mechanismmaxTokens / contextWindow65536 / 196608agent timeout ×2.0trials1 of 6 — 1 pass · 0 fail · 1 errored · 5 pendingmean reward1.00tokens (job total)5,521,115 in / 46,250 outstarted / finished2026-07-08T13:14 / never (aborted)wall clock-Errored trials: RuntimeError (headless-terminal)
| # | result | total | agent | in/out tok | flags | |
|---|---|---|---|---|---|---|
| 1 | ERR | 5m36s | 5m12s | 17516/1165 | 🔍 view |
| # | result | total | agent | in/out tok | flags | |
|---|---|---|---|---|---|---|
| 1 | PASS | 34m14s | 32m26s | 5521115/46250 | 🔍 view |