polyglot-rust-c × 3 at stock settings, budget 8000. This is the suite's canonical hard task (historically failed as a "B1 looper"); the question was whether the bigger budget changes its failure mode.
stop — its tail repeats the thinking
tail verbatim. No tool call → pi's print loop ends "normally" → trial dies
half-done. Crucially this evades the runaway recovery, which only
armed on length stops.pr:2 — two batched advances, vs 79 breaks in the per-call era).The recovery correctly stayed silent on the B2 trials (toolCall present) and had no trigger for FQfZzxS (clean stop). Zero false fires; zero true fires.
harnesses/minimal_pi.py: now also fires
on stopReason "stop" + no toolCall + text > 8000 chars (largest
legitimate final answer observed anywhere today: 1.9k chars; artifact
range 19–75k). Same trim + followUp nudge, same 2/session cap. 16/16 unit
tests incl. three new soft-runaway cases; hysteresis suite still 15/15.Keep 8000: it fixed the artifact where thinking fits (regex-log 6/6 today) and polyglot shows no budget value would fit its thinking anyway — the residual is a harness-recovery problem (now addressed, pending live proof) plus a B2 problem (H2).
llama-local/qwen3.6-35b-a3bagentharnesses.minimal_pi:MinimalPithinkingonreasoning budgetdirect (server/none)* — see journal for the authoritative mechanismmaxTokens / contextWindow32768 / 229376agent timeout ×2.0trials3 of 3 — 0 pass · 2 fail · 1 erroredmean reward0.00tokens (job total)1,766,896 in / 465,591 outstarted / finished2026-07-04T20:52 / 2026-07-04T21:55wall clock1h03mErrored trials: AgentTimeoutError (polyglot-rust-c)
| # | result | total | agent | in/out tok | flags | |
|---|---|---|---|---|---|---|
| 1 | FAIL | 5m20s | 4m24s | 88569/44233 | 🔍 view | |
| 2 | FAIL | 26m45s | 25m43s | 993457/212473 | 🔍 view | |
| 3 | ERR | 30m56s | 30m00s | 684870/208885 | 🔍 view |