First fast-set run of qwen3.5-122b-a10b (MinimalPi, ~5×-scaled cap map
fix-git=600…default=1800 because the stock caps are calibrated on 27b/35b
speed and would cut 122b passes). llama-swap was restarted BEFORE this run
(59 GiB available at start).
Deliberately aborted (SIGINT) at 4/9 after the documented 122b
llama-server memory ratchet (see the model's entry in the llama-swap
config.yaml, kernel OOM kills earlier the same day) burned 59 GiB →
92 MiB available + 11 GiB swap in ~25 minutes of benchmark load.
Extrapolation said swap would exhaust before trial 9, i.e. a near-certain
kernel OOM kill of llama-server mid-run. Aborted at the filter-js trial
boundary; the remaining 5 tasks were rerun as two fresh-restart batches
(fast__…143738, fast__…144656) whose results merge with this run on the
dashboard.
Scored trials (all finished well under their caps, no cut@ markers):
CancelledError at 6m21s / 231k input tokens —
died while the box was thrashing (MemAvailable double-digit MiB); treat as
infrastructure, not model. Clean rerun in fast__…144656 scored FAIL.Combined fast-set verdict across the 3 parts: 4/9 (fix-git, nginx, openssl, sanitize-git-repo) — the 35b MinimalPi baseline pass set minus regex-log (which flaked here but passed in smoke). See COMMENTS.md and the AGENT_TODOS "qwen3.5-122b-a10b operational" item.
💬 3 analyst comments inline below (from runs/fast__qwen3.5-122b-a10b__20260712-141049/COMMENTS.md).
llama-local/qwen3.5-122b-a10bagentharnesses.minimal_pi:MinimalPithinkingonreasoning budgetdirect (server/none)* — see journal for the authoritative mechanismmaxTokens / contextWindow65536 / 196608agent timeout ×2.0trials4 of 9 — 1 pass · 2 fail · 1 errored · 5 pendingmean reward0.25tokens (job total)338,648 in / 20,439 outstarted / finished2026-07-12T14:10 / never (aborted)wall clock-Errored trials: CancelledError (filter-js-from-html)
CancelledError ERR, not a model failure: this trial ran while the
llama-server memory ratchet had the host at double-digit-MiB MemAvailable
and 11 GiB into swap (the run was aborted at this trial's boundary for that
reason). Input hit 231k tokens in 6m21s. The clean rerun in
fast__qwen3.5-122b-a10b__20260712-144656 scored a normal FAIL at 5m37s.
| # | result | total | agent | in/out tok | flags | |
|---|---|---|---|---|---|---|
| 1 | ERR | 6m52s | 6m21s | 231165/8601 | 🔍 view |
| # | result | total | agent | in/out tok | flags | |
|---|---|---|---|---|---|---|
| 1 | PASS | 2m27s | 1m53s | 24488/1602 | 🔍 view |
Clean-stop FAIL at 4m20s, 45.8k input tokens — same shape as the 35b baseline fail (perf bar), nothing 122b-specific spotted at run level.
| # | result | total | agent | in/out tok | flags | |
|---|---|---|---|---|---|---|
| 1 | FAIL | 10m53s | 4m20s | 45816/3359 | 🔍 view |
Failed here at 4m26s agent but PASSED the smoke run 20 minutes earlier at 5m44s on the same model/knobs — a genuine K=1 flake, not a config difference. Worth K=3 before concluding anything about 122b on this task.
| # | result | total | agent | in/out tok | flags | |
|---|---|---|---|---|---|---|
| 1 | FAIL | 5m11s | 4m26s | 37179/6877 | 🔍 view |