The A/B baseline for the RecursivePi read-heavy test
(runs/smoke__qwen3.6-35b-a3b__20260710-130242/NOTES.md). Same model
(qwen3.6-35b-a3b), same current harness, no recursive.
FAIL 0.0, agent 65 s / wall 101 s (faster than both RecursivePi runs, which
paid 40–60 s of wasted child latency). Fails the same two verifier tests as
RecursivePi — test_removal_of_secret_information (a 2nd HF token in a JSON
git-diff string survives) and test_correct_replacement_of_secret_information
(a clean line corrupted / exact-match failure) — while
test_no_other_files_changed passes. No peg-native parser errors this run.
This is the clean isolation: the failure is a model-capability precision problem (faithful transcription of long random secrets, secret-vs-noise judgment, no over-replacement), identical with and without delegation. Recursive neither helps nor is the cause.
llama-local/qwen3.6-35b-a3bagentharnesses.minimal_pi:MinimalPithinkingonreasoning budgetdirect (server/none)* — see journal for the authoritative mechanismmaxTokens / contextWindow65536 / 196608agent timeout ×2.0trials1 of 1 — 0 pass · 1 failmean reward0.00tokens (job total)1,241,890 in / 6,928 outstarted / finished2026-07-10T13:07 / 2026-07-10T13:09wall clock1m40s| # | result | total | agent | in/out tok | flags | |
|---|---|---|---|---|---|---|
| 1 | FAIL | 1m40s | 1m05s | 1241890/6928 | 🔍 view |