← older: fast__qwen3.6-27b__20260708-022105all runs · task boardnewer: fast2__qwen3.6-27b__20260708-051015

fast__qwen3.6-35b-a3b__20260708-034239

fast 35b, tangled 4-knob night batch — COLLAPSE to 3/9; context-strip amnesia is the prime suspect

Night batch 2026-07-08, SubagentsPi, four knobs flipped together vs 07-07: loop_guard=on strip_thinking=on keep_tool_results=10 subagent_nudges=on. NOT a clean A/B (untangling item in AGENT_TODOS.md).

Result 3/9 (mean 0.33) vs 7/9 on the 07-07 SubagentsPi baseline (fast__qwen3.6-35b-a3b__20260707-001626). Regressions: fix-git, cancel-async-tasks, sanitize-git-repo, sparql-university all PASS→FAIL. All regressed trials finished NATURALLY (stopReason=stop, declared done) and the verifier failed them — a quality regression, not timeouts.

Shape of the regression = the amnesia the context-strip ledger item predicted: cancel-async-tasks read /app/run.py 6 times, re-ran near-identical test one-liners in 3–5x clusters, turns ~10→55 (agent 35s→6m19s), then declared done with a ✅-checklist the verifier rejected. Everything got slower (sparql 3m25s→13m25s, sanitize 2m03s→7m02s) and input tokens ballooned (cancel-async 74k→872k — more turns, not bigger turns). Loop guard: 2 blocks total across all 9 roots (quiet, harmless). Contrast: the sibling 27b run with identical knobs went 7/9.

Verdict feeding AGENT_TODOS.md: keep STRIP_THINKING/KEEP_TOOL_RESULTS defaults OFF; rerun 35b with the strip knobs off to confirm they (not loop_guard/nudges) are the regressor.

Run details

modelllama-local/qwen3.6-35b-a3bagentharnesses.subagents_pi:SubagentsPithinkingonreasoning budgetdirect (server/none)* — see journal for the authoritative mechanismmaxTokens / contextWindow65536 / 196608agent timeout ×2.0trials9 of 9 — 3 pass · 6 failmean reward0.33tokens (job total)4,633,120 in / 138,437 outstarted / finished2026-07-08T03:42 / 2026-07-08T05:10wall clock1h27m

Tasks

cancel-async-tasks — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL6m59s6m19s872158/21620
claimed success but the verifier did NOT pass (heuristic)a bash command timed out ×4empty final message (no text, no tool call)loop-guard blocked a repeated call ×2
🔍 view

filter-js-from-html — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL16m54s14m02s500645/10030
🔍 view

fix-git — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL3m38s2m58s338898/6099
claimed success but the verifier did NOT pass (heuristic)
🔍 view

nginx-request-logging — 1/1 passed

#resulttotalagentin/out tokflags
1PASS5m50s5m09s267327/6966
🔍 view

openssl-selfsigned-cert — 1/1 passed

#resulttotalagentin/out tokflags
1PASS5m31s4m53s157352/5704
🔍 view

query-optimize — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL15m05s8m26s290728/10268
a bash command timed outlong reasoning (14,093 chars)
🔍 view

regex-log — 1/1 passed

#resulttotalagentin/out tokflags
1PASS11m36s10m48s441465/36000
long reasoning (21,849 chars) ×2
🔍 view

sanitize-git-repo — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL7m42s7m02s910146/12454
claimed success but the verifier did NOT pass (heuristic)
🔍 view

sparql-university — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL14m13s13m25s854401/29296
claimed success but the verifier did NOT pass (heuristic)long reasoning (19,866 chars)
🔍 view