regex-log × 5 at stock settings, first K>1 validation of the reasoning-budget switch 4000 → 8000 (~19:00 same day). Direct comparison against …180910 (same task × 5 at budget 4000: 3/5, both fails were force-close → plain-text thinking-continuation runaways).
</think> and the model comes
out clean every turn.This is the cleanest possible confirmation of the force-close hypothesis: the
task whose loops motivated the entire reasoning-budget mechanism goes 6/6
today (counting …192605) once the budget exceeds its natural thinking length.
The artifact was never about the task being hard — it was the mid-thought
</think> injection.
Caveat from the sibling run …205241: 8000 only helps tasks whose thinking fits. polyglot-rust-c maxes any budget and still produces the (milder, self-terminating) continuation artifact. Keep 8000; attack the residual with the extended runaway recovery, not with more budget.
llama-local/qwen3.6-35b-a3bagentharnesses.minimal_pi:MinimalPithinkingonreasoning budgetdirect (server/none)* — see journal for the authoritative mechanismmaxTokens / contextWindow32768 / 229376agent timeout ×2.0trials5 of 5 — 5 pass · 0 failmean reward1.00tokens (job total)1,138,589 in / 109,498 outstarted / finished2026-07-04T20:36 / 2026-07-04T20:52wall clock15m52s| # | result | total | agent | in/out tok | flags | |
|---|---|---|---|---|---|---|
| 1 | PASS | 2m42s | 1m45s | 176142/17363 | 🔍 view | |
| 2 | PASS | 3m22s | 2m24s | 224952/23397 | 🔍 view | |
| 3 | PASS | 4m02s | 3m03s | 301145/29978 | 🔍 view | |
| 4 | PASS | 3m13s | 2m16s | 304281/23688 | 🔍 view | |
| 5 | PASS | 2m31s | 1m32s | 132069/15072 | 🔍 view |