First-ever benchmark run of the 122B MoE (10B active, CPU-offloaded experts,
-c 262144, DRY removed server-side 2026-07-12). MinimalPi pinned, all knobs
current defaults, uncapped.
Result 4/4 PASS (fix-git 1m21s, nginx 1m35s, openssl 1m54s, regex-log 5m44s agent). Zero schema/tool errors, no wire-format leaks, no loop-guard or write-guard events — the qwen tool surface + XML template worked cleanly through llama.cpp.
Speed calibration: short tasks run ~4–7× the 35b's agent time (fix-git 13s→81s,
nginx 14s→95s, openssl 16-30s→114s), regex-log ~1.7× (3m20s→5m44s). This is
why the follow-up fast run used a ~5×-scaled cap map — the stock
TASK_SET_fast_TIMEOUTS would have cut openssl 6 seconds short of its pass.
Operational finding (bit hard in the fast run, see
fast__qwen3.5-122b-a10b__20260712-141049/NOTES.md): the 122b llama-server
private-memory ratchet consumed ~5 GiB of RAM during just this 13.5-minute
run (MemAvailable 5.3 GiB → 328 MiB).
llama-local/qwen3.5-122b-a10bagentharnesses.minimal_pi:MinimalPithinkingonreasoning budgetdirect (server/none)* — see journal for the authoritative mechanismmaxTokens / contextWindow65536 / 196608agent timeout ×2.0trials4 of 4 — 4 pass · 0 failmean reward1.00tokens (job total)115,408 in / 14,685 outstarted / finished2026-07-12T13:55 / 2026-07-12T14:09wall clock13m34s| # | result | total | agent | in/out tok | flags | |
|---|---|---|---|---|---|---|
| 1 | PASS | 1m55s | 1m21s | 24903/1709 | 🔍 view |
| # | result | total | agent | in/out tok | flags | |
|---|---|---|---|---|---|---|
| 1 | PASS | 2m40s | 1m35s | 37616/2308 | 🔍 view |
| # | result | total | agent | in/out tok | flags | |
|---|---|---|---|---|---|---|
| 1 | PASS | 2m29s | 1m54s | 32278/2225 | 🔍 view |
| # | result | total | agent | in/out tok | flags | |
|---|---|---|---|---|---|---|
| 1 | PASS | 6m29s | 5m44s | 20611/8443 | 🔍 view |