← older: smoke__qwen3.5-122b-a10b__20260712-135546all runs · task boardnewer: fast__qwen3.5-122b-a10b__20260712-143738

fast__qwen3.5-122b-a10b__20260712-141049

122b fast run part 1/3 — ABORTED 4/9 at a trial boundary: llama-server RAM ratchet nearly OOMed the box

First fast-set run of qwen3.5-122b-a10b (MinimalPi, ~5×-scaled cap map fix-git=600…default=1800 because the stock caps are calibrated on 27b/35b speed and would cut 122b passes). llama-swap was restarted BEFORE this run (59 GiB available at start).

Deliberately aborted (SIGINT) at 4/9 after the documented 122b llama-server memory ratchet (see the model's entry in the llama-swap config.yaml, kernel OOM kills earlier the same day) burned 59 GiB → 92 MiB available + 11 GiB swap in ~25 minutes of benchmark load. Extrapolation said swap would exhaust before trial 9, i.e. a near-certain kernel OOM kill of llama-server mid-run. Aborted at the filter-js trial boundary; the remaining 5 tasks were rerun as two fresh-restart batches (fast__…143738, fast__…144656) whose results merge with this run on the dashboard.

Scored trials (all finished well under their caps, no cut@ markers):

Combined fast-set verdict across the 3 parts: 4/9 (fix-git, nginx, openssl, sanitize-git-repo) — the 35b MinimalPi baseline pass set minus regex-log (which flaked here but passed in smoke). See COMMENTS.md and the AGENT_TODOS "qwen3.5-122b-a10b operational" item.

💬 3 analyst comments inline below (from runs/fast__qwen3.5-122b-a10b__20260712-141049/COMMENTS.md).

Run details

modelllama-local/qwen3.5-122b-a10bagentharnesses.minimal_pi:MinimalPithinkingonreasoning budgetdirect (server/none)* — see journal for the authoritative mechanismmaxTokens / contextWindow65536 / 196608agent timeout ×2.0trials4 of 9 — 1 pass · 2 fail · 1 errored · 5 pendingmean reward0.25tokens (job total)338,648 in / 20,439 outstarted / finished2026-07-12T14:10 / never (aborted)wall clock-

Errored trials: CancelledError (filter-js-from-html)

Tasks

filter-js-from-html — 0/1 passed

💬 analyst comment

CancelledError ERR, not a model failure: this trial ran while the llama-server memory ratchet had the host at double-digit-MiB MemAvailable and 11 GiB into swap (the run was aborted at this trial's boundary for that reason). Input hit 231k tokens in 6m21s. The clean rerun in fast__qwen3.5-122b-a10b__20260712-144656 scored a normal FAIL at 5m37s.

#resulttotalagentin/out tokflags
1ERR6m52s6m21s231165/8601
trial errored: CancelledError
🔍 view

fix-git — 1/1 passed

#resulttotalagentin/out tokflags
1PASS2m27s1m53s24488/1602
🔍 view

query-optimize — 0/1 passed

💬 analyst comment

Clean-stop FAIL at 4m20s, 45.8k input tokens — same shape as the 35b baseline fail (perf bar), nothing 122b-specific spotted at run level.

#resulttotalagentin/out tokflags
1FAIL10m53s4m20s45816/3359
a bash command timed out
🔍 view

regex-log — 0/1 passed

💬 analyst comment

Failed here at 4m26s agent but PASSED the smoke run 20 minutes earlier at 5m44s on the same model/knobs — a genuine K=1 flake, not a config difference. Worth K=3 before concluding anything about 122b on this task.

#resulttotalagentin/out tokflags
1FAIL5m11s4m26s37179/6877
long reasoning (13,527 chars)
🔍 view