← older: medic__qwen3.6-35b-a3b__20260710-232538all runs · task boardnewer: fast__qwen3.5-122b-a10b__20260712-141049

smoke__qwen3.5-122b-a10b__20260712-135546

qwen3.5-122b-a10b first contact — smoke 4/4, harness works, ~4-7× slower than 35b

First-ever benchmark run of the 122B MoE (10B active, CPU-offloaded experts, -c 262144, DRY removed server-side 2026-07-12). MinimalPi pinned, all knobs current defaults, uncapped.

Result 4/4 PASS (fix-git 1m21s, nginx 1m35s, openssl 1m54s, regex-log 5m44s agent). Zero schema/tool errors, no wire-format leaks, no loop-guard or write-guard events — the qwen tool surface + XML template worked cleanly through llama.cpp.

Speed calibration: short tasks run ~4–7× the 35b's agent time (fix-git 13s→81s, nginx 14s→95s, openssl 16-30s→114s), regex-log ~1.7× (3m20s→5m44s). This is why the follow-up fast run used a ~5×-scaled cap map — the stock TASK_SET_fast_TIMEOUTS would have cut openssl 6 seconds short of its pass.

Operational finding (bit hard in the fast run, see fast__qwen3.5-122b-a10b__20260712-141049/NOTES.md): the 122b llama-server private-memory ratchet consumed ~5 GiB of RAM during just this 13.5-minute run (MemAvailable 5.3 GiB → 328 MiB).

Run details

modelllama-local/qwen3.5-122b-a10bagentharnesses.minimal_pi:MinimalPithinkingonreasoning budgetdirect (server/none)* — see journal for the authoritative mechanismmaxTokens / contextWindow65536 / 196608agent timeout ×2.0trials4 of 4 — 4 pass · 0 failmean reward1.00tokens (job total)115,408 in / 14,685 outstarted / finished2026-07-12T13:55 / 2026-07-12T14:09wall clock13m34s

Tasks

fix-git — 1/1 passed

#resulttotalagentin/out tokflags
1PASS1m55s1m21s24903/1709
🔍 view

nginx-request-logging — 1/1 passed

#resulttotalagentin/out tokflags
1PASS2m40s1m35s37616/2308
🔍 view

openssl-selfsigned-cert — 1/1 passed

#resulttotalagentin/out tokflags
1PASS2m29s1m54s32278/2225
🔍 view

regex-log — 1/1 passed

#resulttotalagentin/out tokflags
1PASS6m29s5m44s20611/8443
long reasoning (26,095 chars)
🔍 view