← older: fast2__qwen3.5-122b-a10b__20260712-173102all runs · task boardnewer: fast__qwen3.5-122b-a10b__20260712-183411

fast2__qwen3.5-122b-a10b__20260712-180346

122b fast2 7/7 — headless-terminal rerun (cap 1700): FAIL, scored cleanly; sequence total 2/6

Rerun of the cap-collision ERR in fast2__qwen3.5-122b-a10b__20260712-162421 (cap 1700s < the 1800s declared×mult ceiling this time). headless-terminal FAIL at 10m36s agent, 387k input tokens — finished and declared done on its own (no cut), unlike the first attempt's full-30m grind; the verifier rejected the result. Same verdict as the 35b, now properly scored.

122b fast2 sequence verdict (7 jobs, 6 tasks): 2/6git-multibranch PASS (8m08s) and mailman PASS (22m49s) are the first fast2 passes ever recorded on this box (35b/27b: 0/6 in every prior run); reshard-c4-data FAIL 12m37s, headless-terminal FAIL 10m36s (this rerun), financial-document-processor FAIL cut@30m, fix-ocaml-gc FAIL cut@30m. Combined with the fast set (4/9, pass set = 35b's minus a regex-log flake), the 122b's edge shows exactly where predicted: multi-step tasks the smaller model can't hold together, not the hidden-criterion near-misses.

Sequence infrastructure notes: one task per job; the between-jobs llama-swap auto-restart (shipped this session) fired before all 7 jobs; swap watchdog never tripped; caps ×3 stock clamped at 1800s (mailman's pass at 1369s would have been cut by the stock 600s cap).

Run details

modelllama-local/qwen3.5-122b-a10bagentharnesses.minimal_pi:MinimalPithinkingonreasoning budgetdirect (server/none)* — see journal for the authoritative mechanismmaxTokens / contextWindow65536 / 196608agent timeout ×2.0trials1 of 1 — 0 pass · 1 failmean reward0.00tokens (job total)386,993 in / 11,726 outstarted / finished2026-07-12T18:03 / 2026-07-12T18:15wall clock11m46s

Tasks

headless-terminal — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL11m46s10m36s386993/11726
🔍 view