← older: smoke__qwen3.6-35b-a3b__20260704-193517all runs · task boardnewer: smoke__qwen3.6-35b-a3b__20260704-205241

smoke__qwen3.6-35b-a3b__20260704-203646

regex-log × 5 at budget 8000 — 5/5, force-close artifact GONE on this task

What was run

regex-log × 5 at stock settings, first K>1 validation of the reasoning-budget switch 4000 → 8000 (~19:00 same day). Direct comparison against …180910 (same task × 5 at budget 4000: 3/5, both fails were force-close → plain-text thinking-continuation runaways).

Results

Interpretation

This is the cleanest possible confirmation of the force-close hypothesis: the task whose loops motivated the entire reasoning-budget mechanism goes 6/6 today (counting …192605) once the budget exceeds its natural thinking length. The artifact was never about the task being hard — it was the mid-thought </think> injection.

Caveat from the sibling run …205241: 8000 only helps tasks whose thinking fits. polyglot-rust-c maxes any budget and still produces the (milder, self-terminating) continuation artifact. Keep 8000; attack the residual with the extended runaway recovery, not with more budget.

Run details

modelllama-local/qwen3.6-35b-a3bagentharnesses.minimal_pi:MinimalPithinkingonreasoning budgetdirect (server/none)* — see journal for the authoritative mechanismmaxTokens / contextWindow32768 / 229376agent timeout ×2.0trials5 of 5 — 5 pass · 0 failmean reward1.00tokens (job total)1,138,589 in / 109,498 outstarted / finished2026-07-04T20:36 / 2026-07-04T20:52wall clock15m52s

Tasks

regex-log — 5/5 passed

#resulttotalagentin/out tokflags
1PASS2m42s1m45s176142/17363
long reasoning (22,130 chars)
🔍 view
2PASS3m22s2m24s224952/23397
long reasoning (20,566 chars)
🔍 view
3PASS4m02s3m03s301145/29978
long reasoning (22,210 chars) ×2
🔍 view
4PASS3m13s2m16s304281/23688
long reasoning (20,655 chars)
🔍 view
5PASS2m31s1m32s132069/15072
long reasoning (20,135 chars)
🔍 view