← older: fast__qwen3.5-122b-a10b__20260712-143738all runs · task boardnewer: fast2__qwen3.5-122b-a10b__20260712-153513

fast__qwen3.5-122b-a10b__20260712-144656

122b fast run part 3/3 — final batch: sanitize-git-repo PASS, sparql + filter-js FAIL; combined fast set 4/9

Final resume batch of the memory-aborted fast__…141049 (MinimalPi, scaled caps, llama-swap restarted before the batch).

This batch rode the ratchet to 29 MiB MemAvailable / 13 GiB swap near the end but finished without OOM (bounded by the last trial's 900s cap; swap never dropped below ~6.7 GiB free).

Combined 122b fast-set verdict (parts 1–3): 4/9 — PASS fix-git, nginx-request-logging, openssl-selfsigned-cert, sanitize-git-repo. That is the 35b MinimalPi baseline pass set (5/9) minus regex-log, which flaked (smoke PASS, fast FAIL). At ~4–7× the wall clock per task, the 122b shows no fast-set advantage over the 35b so far; the smaller model's failures here are hidden-criterion/task-logic, and more parameters didn't buy those back. No trial hit its cap, so the scores are cap-clean.

Run details

modelllama-local/qwen3.5-122b-a10bagentharnesses.minimal_pi:MinimalPithinkingonreasoning budgetdirect (server/none)* — see journal for the authoritative mechanismmaxTokens / contextWindow65536 / 196608agent timeout ×2.0trials3 of 3 — 1 pass · 2 failmean reward0.33tokens (job total)862,090 in / 21,670 outstarted / finished2026-07-12T14:46 / 2026-07-12T15:10wall clock23m08s

Tasks

filter-js-from-html — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL8m48s5m37s141514/7576
🔍 view

sanitize-git-repo — 1/1 passed

#resulttotalagentin/out tokflags
1PASS7m38s6m56s584477/6059
🔍 view

sparql-university — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL6m40s5m56s136099/8035
long reasoning (18,984 chars)
🔍 view