Job 2 of the four one-task first-light jobs. MinimalPi, K=1, uncapped.
Batch roll-up: smoke__laguna-s-2.1__20260726-201248/NOTES.md.
The long one, and the most informative pass of the batch: 21m30s of agent time for 24778 output tokens ≈ 19 tok/s sustained, consistent with the ~21 tok/s measured on short probes — i.e. no degradation across a real multi-turn session with a 225k-token cumulative prompt. No length-stops, no guard fires.
Two things worth carrying forward:
TASK_SET_fast_TIMEOUTS
caps regex-log at 600s; laguna needed 1290s of agent time to pass it. Running
the fast set with stock caps would score this a FAIL that isn't one.smoke__qwen3.6-35b-a3b__20260726-100656, so don't read this as laguna
beating the 35b on capability — it's a slow-but-completes vs cut-at-cap
comparison, not a clean one. The 35b passes regex-log uncapped historically.llama-local/laguna-s-2.1agentharnesses.minimal_pi:MinimalPithinkingonreasoning budgetdirect (server/none)* — see journal for the authoritative mechanismmaxTokens / contextWindow65536 / 196608agent timeout ×2.0trials1 of 1 — 1 pass · 0 failmean reward1.00tokens (job total)225,315 in / 24,778 outstarted / finished2026-07-26T19:42 / 2026-07-26T20:04wall clock22m16s| # | result | total | agent | in/out tok | flags | |
|---|---|---|---|---|---|---|
| 1 | PASS | 22m16s | 21m30s | 225315/24778 | 🔍 view |