fix-ocaml-gc FAIL cut@1800s (same as MinimalPi; zero delegation, no loops). Sequence verdict (8 never-passed tasks, SubagentsPi, one job each): 3 flips — reshard-c4-data 23m10s, headless-terminal 19m38s, financial-document-processor 12m25s — ALL on fast2 multi-step tasks; 0/4 on the fast hidden-criterion tasks (cancel-async/sparql/filter-js/ query-optimize), which fail on unstated bars no workflow reaches.
MECHANISM CAVEAT (checked per-trial): none of the 3 passes used scout/planner; 2 used only the FORCED reviewer, which verified without changing anything; headless passed with zero children. The causal candidates are the staged workflow PROMPT (roots visibly persisted longer / worked more systematically) and K=1 variance — NOT delegation. Ledger items: K=3 replication A/B (workflow prompt vs plain MinimalPi), and the lean 122b profile (prompt + force_review only).
122b cumulative after today: fast2 5/6 solved by some 122b config (git-multibranch + mailman MinimalPi; reshard + headless + financial SubagentsPi; only fix-ocaml-gc unsolved), fast still 4/9 + the regex-log smoke flake.
llama-local/qwen3.5-122b-a10bagentharnesses.subagents_pi:SubagentsPithinkingonreasoning budgetdirect (server/none)* — see journal for the authoritative mechanismmaxTokens / contextWindow65536 / 196608agent timeout ×2.0trials1 of 1 — 0 pass · 1 failmean reward0.00tokens (job total)1,501,245 in / 34,327 outstarted / finished2026-07-12T20:56 / 2026-07-12T21:29wall clock32m17s| # | result | total | agent | in/out tok | flags | |
|---|---|---|---|---|---|---|
| 1 | FAIL | 32m17s | 30m00s | 1501245/34327 | 🔍 view |