← older: smoke__qwen3.6-35b-a3b__20260704-180910all runs · task boardnewer: smoke__qwen3.6-35b-a3b__20260704-193517

smoke__qwen3.6-35b-a3b__20260704-192605

First budget-8000 smoke — 3/4, flag confirmed in effect, openssl fail is baseline

What was run

fix-git + nginx-request-logging + openssl-selfsigned-cert + regex-log × 1, first run after the server-side --reasoning-budget switch 4000 → 8000 (~19:00, both qwen bases in the llama-swap config; llama-swap hot-reloaded).

Results

Verdict

Nothing unexpected. The interesting 8000 evidence is in the follow-up runs: regex-log × 5 (…203646, 5/5) and polyglot-rust-c × 3 (…205241, artifact survives in soft form on thinking-heavy tasks).

Run details

modelllama-local/qwen3.6-35b-a3bagentharnesses.minimal_pi:MinimalPithinkingonreasoning budgetdirect (server/none)* — see journal for the authoritative mechanismmaxTokens / contextWindow32768 / 229376agent timeout ×2.0trials4 of 4 — 3 pass · 1 failmean reward0.75tokens (job total)858,538 in / 53,491 outstarted / finished2026-07-04T19:26 / 2026-07-04T19:35wall clock9m10s

Tasks

fix-git — 1/1 passed

#resulttotalagentin/out tokflags
1PASS56s9s20865/1312
🔍 view

nginx-request-logging — 1/1 passed

#resulttotalagentin/out tokflags
1PASS1m04s14s49764/2157
🔍 view

openssl-selfsigned-cert — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL1m06s18s53006/2591
claimed success but the verifier did NOT pass (heuristic)
🔍 view

regex-log — 1/1 passed

#resulttotalagentin/out tokflags
1PASS6m02s5m04s734903/47431
long reasoning (20,511 chars) ×2
🔍 view