← older: fast2__qwen3.5-122b-a10b__20260712-205638all runs · task boardnewer: suite__qwen3.6-35b-a3b__20260720-025830

suite__qwen3.6-35b-a3b__20260718-104156

Full suite, post-fix checkpoint (MinimalPi, K=1)

35/89 = 39.3% (baseline suite 20260703: 33/89 = 37.1%). Elapsed 26h45m. Purpose: user-requested checkpoint of the whole fix stack (pi 0.80.2 + print-mode/exit patches, maxTokens 65536, qwen-tools surface, preamble, write/loop guards, bash timeout 90s, reasoning-budget 8000 server-side) before deciding on an official K=5 leaderboard run. Config: all config/harness.env defaults, AGENT_TIMEOUT_MULT=2.0, no fast-fail caps.

Headline

Leaderboard context

Comfortably above little-coder (24.6 ± 3.2 on the official board). Official submission would target terminal-bench 2.1 (2.0 board closed), require AGENT_TIMEOUT_MULT=1.0 (this run used 2.0 — expect some of the long passes to become fails there) and K=5.

💬 8 analyst comments inline below (from runs/suite__qwen3.6-35b-a3b__20260718-104156/COMMENTS.md).

Run details

modelllama-local/qwen3.6-35b-a3bagentharnesses.minimal_pi:MinimalPithinkingonreasoning budgetdirect (server/none)* — see journal for the authoritative mechanismmaxTokens / contextWindow65536 / 196608agent timeout ×2.0trials89 of 89 — 35 pass · 42 fail · 12 erroredmean reward0.39tokens (job total)533,689,992 in / 7,225,142 outstarted / finished2026-07-18T10:41 / 2026-07-19T13:27wall clock26h45m

Errored trials: AgentTimeoutError (caffe-cifar-10, circuit-fibsqrt, crack-7z-hash, feal-linear-cryptanalysis, make-mips-interpreter, qemu-alpine-ssh, qemu-startup, rstan-to-pystan, torch-pipeline-parallelism, train-fasttext, write-compressor) · NonZeroAgentExitCodeError (pytorch-model-recovery)

Tasks

adaptive-rejection-sampler — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL21m17s19m26s10160149/81162
🔍 view

bn-fit-modify — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL8m58s8m09s3364723/68210
claimed success but the verifier did NOT pass (heuristic)long reasoning (22,403 chars) ×4
🔍 view

break-filter-js-from-html — 1/1 passed

#resulttotalagentin/out tokflags
1PASS6m28s5m53s595199/48163
long reasoning (29,459 chars) ×6
🔍 view

build-cython-ext — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL2m35s2m02s865842/9928
🔍 view

build-pmars — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL1m34s58s372571/5372
🔍 view

build-pov-ray — 1/1 passed

#resulttotalagentin/out tokflags
1PASS9m04s6m19s3009990/16665
a bash command timed out ×2
🔍 view

caffe-cifar-10 — 0/1 passed

#resulttotalagentin/out tokflags
1ERR40m44s40m00s1979199/13602
trial errored: AgentTimeoutErrora bash command timed out ×4
🔍 view

cancel-async-tasks — 0/1 passed

💬 analyst comment

New fast-false-done shape for this task: wrote run.py, ran a 2-command self-test, declared "Done." at 19s agent time. Its own check passed; the hidden criteria did not. Different from the earlier deadlocked-smoke-test shape (fast__…113104) — this is the hidden-criterion near-miss class, K≥3 territory.

#resulttotalagentin/out tokflags
1FAIL1m05s19s17146/2894
🔍 view

chess-best-move — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL13m38s12m58s4327775/106906
long reasoning (25,580 chars) ×6
🔍 view

circuit-fibsqrt — 0/1 passed

#resulttotalagentin/out tokflags
1ERR2h00m2h00m24915750/879301
trial errored: AgentTimeoutErrorlong reasoning (26,794 chars) ×7loop-guard blocked a repeated call ×2
🔍 view

cobol-modernization — 1/1 passed

#resulttotalagentin/out tokflags
1PASS2m25s1m51s687600/17966
long reasoning (12,693 chars)
🔍 view

code-from-image — 1/1 passed

#resulttotalagentin/out tokflags
1PASS3m23s2m49s152375/10786
a bash command timed out
🔍 view

compile-compcert — 1/1 passed

#resulttotalagentin/out tokflags
1PASS22m18s21m33s455660/5486
a bash command timed out
🔍 view

configure-git-webserver — 1/1 passed

#resulttotalagentin/out tokflags
1PASS1m04s22s38391/2414
🔍 view

constraints-scheduling — 1/1 passed

#resulttotalagentin/out tokflags
1PASS1m06s24s21538/4249
🔍 view

count-dataset-tokens — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL6m46s6m13s3234931/20201
🔍 view

crack-7z-hash — 0/1 passed

#resulttotalagentin/out tokflags
1ERR31m56s30m00s4059526/70608
claimed success but the verifier did NOT pass (heuristic)trial errored: AgentTimeoutErrora bash command timed out ×4
🔍 view

custom-memory-heap-crash — 1/1 passed

#resulttotalagentin/out tokflags
1PASS2m54s2m09s889523/16754
long reasoning (21,955 chars)
🔍 view

db-wal-recovery — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL10m40s9m59s1432092/79392
long reasoning (25,326 chars) ×5
🔍 view

distribution-search — 1/1 passed

#resulttotalagentin/out tokflags
1PASS9m41s9m05s1829459/80799
long reasoning (17,433 chars) ×6
🔍 view

dna-assembly — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL27m12s26m28s9763837/201341
long reasoning (24,977 chars) ×10
🔍 view

dna-insert — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL22m50s22m05s27580359/156237
long reasoning (19,591 chars) ×2loop-guard blocked a repeated call
🔍 view

extract-elf — 1/1 passed

#resulttotalagentin/out tokflags
1PASS7m00s6m20s1801087/58209
long reasoning (19,299 chars)
🔍 view

extract-moves-from-video — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL24m39s23m56s1513840/37966
a bash command timed outlong reasoning (26,889 chars)
🔍 view

feal-differential-cryptanalysis — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL51m18s50m31s8881353/248476
loop-guard hard-stopped the sessionempty final message (no text, no tool call)long reasoning (20,667 chars) ×7loop-guard blocked a repeated call ×13loop-guard escalation nudge
🔍 view

feal-linear-cryptanalysis — 0/1 passed

💬 analyst comment

Baseline-PASS → ERR regression checked: no guard interference (zero loop-guard blocks; one legitimate empty-final recovery nudge late in the session). The model spent the 1h cap rewriting a multithreaded brute-force attack in C — hard-task variance, not harness.

#resulttotalagentin/out tokflags
1ERR1h04m1h00m4406467/292016
trial errored: AgentTimeoutErrora bash command timed out ×2empty final message (no text, no tool call)long reasoning (22,118 chars) ×17runaway / empty-final recovery fired
🔍 view

filter-js-from-html — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL4m14s55s248203/9184
🔍 view

financial-document-processor — 1/1 passed

#resulttotalagentin/out tokflags
1PASS4m54s4m09s1576444/36180
🔍 view

fix-code-vulnerability — 1/1 passed

#resulttotalagentin/out tokflags
1PASS1m17s45s520223/4302
🔍 view

fix-git — 1/1 passed

#resulttotalagentin/out tokflags
1PASS51s16s47387/2336
🔍 view

fix-ocaml-gc — 1/1 passed

#resulttotalagentin/out tokflags
1PASS34m38s27m50s24528341/147395
long reasoning (21,184 chars) ×14runaway / empty-final recovery fired ×2
🔍 view

gcode-to-text — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL2m43s2m09s1447417/14845
🔍 view

git-leak-recovery — 1/1 passed

#resulttotalagentin/out tokflags
1PASS50s12s21359/1621
🔍 view

git-multibranch — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL3m21s2m37s387290/10075
🔍 view

gpt2-codegolf — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL23m04s22m23s8450921/187284
long reasoning (24,825 chars) ×11
🔍 view

headless-terminal — 1/1 passed

#resulttotalagentin/out tokflags
1PASS2m14s1m24s143099/9241
🔍 view

hf-model-inference — 1/1 passed

#resulttotalagentin/out tokflags
1PASS3m53s3m19s307132/10530
🔍 view

install-windows-3.11 — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL6m01s5m09s1471716/19529
a bash command timed out
🔍 view

kv-store-grpc — 1/1 passed

#resulttotalagentin/out tokflags
1PASS52s20s48539/2260
🔍 view

large-scale-text-editing — 1/1 passed

#resulttotalagentin/out tokflags
1PASS13m09s12m03s4953654/80405
a bash command timed out ×3long reasoning (22,478 chars) ×8
🔍 view

largest-eigenval — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL12m19s11m47s3361619/99761
long reasoning (12,813 chars) ×7
🔍 view

llm-inference-batching-scheduler — 1/1 passed

#resulttotalagentin/out tokflags
1PASS11m48s11m13s4454812/105796
long reasoning (21,106 chars) ×7
🔍 view

log-summary-date-ranges — 1/1 passed

#resulttotalagentin/out tokflags
1PASS1m03s29s88121/4772
🔍 view

mailman — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL5m44s4m23s2003681/20811
claimed success but the verifier did NOT pass (heuristic)
🔍 view

make-doom-for-mips — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL24m05s22m55s42205910/120978
🔍 view

make-mips-interpreter — 0/1 passed

#resulttotalagentin/out tokflags
1ERR1h01m1h00m13916675/245418
trial errored: AgentTimeoutErrorlong reasoning (17,278 chars) ×7
🔍 view

mcmc-sampling-stan — 1/1 passed

#resulttotalagentin/out tokflags
1PASS34m03s19m02s151146/8508
🔍 view

merge-diff-arc-agi-task — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL5m26s4m03s533106/40512
long reasoning (16,315 chars) ×3
🔍 view

model-extraction-relu-logits — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL19m07s18m26s10014842/135772
long reasoning (24,547 chars) ×9
🔍 view

modernize-scientific-stack — 1/1 passed

#resulttotalagentin/out tokflags
1PASS1m00s17s37215/2277
🔍 view

mteb-leaderboard — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL7m00s6m22s3218806/22009
a bash command timed out
🔍 view

mteb-retrieve — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL15m32s10m33s1046389/84924
claimed success but the verifier did NOT pass (heuristic)long reasoning (16,667 chars) ×4runaway / empty-final recovery fired
🔍 view

multi-source-data-merger — 1/1 passed

#resulttotalagentin/out tokflags
1PASS1m05s22s25816/3353
🔍 view

nginx-request-logging — 1/1 passed

#resulttotalagentin/out tokflags
1PASS52s17s41821/2398
🔍 view

openssl-selfsigned-cert — 1/1 passed

#resulttotalagentin/out tokflags
1PASS59s25s50569/4167
🔍 view

overfull-hbox — 1/1 passed

#resulttotalagentin/out tokflags
1PASS17m12s16m15s10448137/136846
long reasoning (22,577 chars) ×6loop-guard blocked a repeated call ×2
🔍 view

password-recovery — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL17m39s16m58s5400630/140512
loop-guard hard-stopped the sessionempty final message (no text, no tool call)long reasoning (21,435 chars) ×4loop-guard blocked a repeated call ×10loop-guard escalation nudge
🔍 view

path-tracing — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL6m47s6m06s1286297/43720
long reasoning (17,982 chars) ×2
🔍 view

path-tracing-reverse — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL48m37s47m51s24434840/274778
loop-guard hard-stopped the sessionlong reasoning (23,693 chars) ×10loop-guard blocked a repeated call ×12loop-guard escalation nudge
🔍 view

polyglot-c-py — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL7m39s6m56s910803/64057
claimed success but the verifier did NOT pass (heuristic)long reasoning (25,546 chars) ×7
🔍 view

polyglot-rust-c — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL4m20s3m07s47614/33033
empty final message (no text, no tool call)long reasoning (24,782 chars) ×4runaway / empty-final recovery fired
🔍 view

portfolio-optimization — 1/1 passed

#resulttotalagentin/out tokflags
1PASS2m13s50s30019/2162
🔍 view

protein-assembly — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL18m53s18m17s4886647/106556
long reasoning (25,404 chars) ×7
🔍 view

prove-plus-comm — 1/1 passed

#resulttotalagentin/out tokflags
1PASS1m39s57s175351/8962
🔍 view

pypi-server — 1/1 passed

#resulttotalagentin/out tokflags
1PASS59s25s68688/2156
🔍 view

pytorch-model-cli — 1/1 passed

#resulttotalagentin/out tokflags
1PASS8m48s7m06s1035499/20762
🔍 view

pytorch-model-recovery — 0/1 passed

💬 analyst comment

Instant death, 0 tokens: pi rejected the task prompt itself — Error: Unknown option: - You are given a PyTorch state dictionary…. instruction.md is the only one of the 89 that starts with -; pi's CLI (cli/args.js) errors on any positional starting with a single dash and has no -- terminator. The 20260703 baseline trial died identically, so this task has NEVER run in any suite. Harness fix: prefix dash-leading instructions in MinimalPi.run().

#resulttotalagentin/out tokflags
1ERR4m56s0s0/0
trial errored: NonZeroAgentExitCodeError
🔍 view

qemu-alpine-ssh — 0/1 passed

#resulttotalagentin/out tokflags
1ERR30m36s30m00s1663724/30228
trial errored: AgentTimeoutErrora bash command timed out ×11
🔍 view

qemu-startup — 0/1 passed

💬 analyst comment

Baseline-PASS → ERR regression checked: no guard interference. 100 shell commands / 39 write_file thrashing QEMU serial-over-telnet bridging approaches (-serial pty + socat at the end) until the 30-min cap. NB the 20260703 "pass" was itself a pass-despite-AgentTimeout, so this task was always marginal at this budget.

#resulttotalagentin/out tokflags
1ERR32m00s30m00s6679844/153197
trial errored: AgentTimeoutErrora bash command timed outlong reasoning (27,477 chars) ×5
🔍 view

query-optimize — 1/1 passed

#resulttotalagentin/out tokflags
1PASS10m24s3m53s70804/3110
a bash command timed out
🔍 view

raman-fitting — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL6m37s6m04s1931230/53910
🔍 view

regex-chess — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL1h15m1h14m45076293/481050
loop-guard hard-stopped the sessionlong reasoning (34,665 chars) ×17loop-guard blocked a repeated call ×12loop-guard escalation nudge
🔍 view

regex-log — 1/1 passed

#resulttotalagentin/out tokflags
1PASS21m01s20m17s5614010/172962
long reasoning (18,918 chars) ×2
🔍 view

reshard-c4-data — 1/1 passed

#resulttotalagentin/out tokflags
1PASS3m50s2m15s320565/12819
🔍 view

rstan-to-pystan — 0/1 passed

#resulttotalagentin/out tokflags
1ERR1h00m1h00m1952562/16211
trial errored: AgentTimeoutErrora bash command timed out ×2
🔍 view

sam-cell-seg — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL12m04s10m32s2776397/54606
long reasoning (26,766 chars)
🔍 view

sanitize-git-repo — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL5m20s4m47s5000652/26543
claimed success but the verifier did NOT pass (heuristic)
🔍 view

schemelike-metacircular-eval — 0/1 passed

💬 analyst comment

765 assistant turns / 105.8M input tokens in 66 min (~20% of the entire run's input) grinding ~5s turns at ~170k ctx, then self-declared done and failed. Nothing bounds turn count or per-trial input tokens — recorded as a wall-clock-lever observation in the ledger.

#resulttotalagentin/out tokflags
1FAIL1h07m1h06m105783730/365299
long reasoning (30,860 chars) ×14loop-guard blocked a repeated call ×5
🔍 view

sparql-university — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL2m44s1m59s105595/19925
long reasoning (24,783 chars) ×2
🔍 view

sqlite-db-truncate — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL15m35s15m01s2166749/136345
claimed success but the verifier did NOT pass (heuristic)long reasoning (16,981 chars) ×8runaway / empty-final recovery fired
🔍 view

sqlite-with-gcov — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL1m26s42s73623/2408
🔍 view

torch-pipeline-parallelism — 0/1 passed

💬 analyst comment

wg:2547 is not 2547 generations — it is ONE garbled length-stopped generation. The model looped emitting Qwen <tool_call> XML as text; the parser split it into ~2547 near-identical write_file calls with wire markup leaked inside the content argument. The write guard blocked every one (correct per-call), but pi processed the whole queue until the 30-min AgentTimeout killed it mid-write of the giant message_end line (truncated at exactly the 64KiB pipe buffer). No escalation path exists for this shape — the loop guard never counts write-guard blocks. Proposed per-message blocked-call cap in the ledger.

#resulttotalagentin/out tokflags
1ERR36m04s30m00s1026758/45248
trial errored: AgentTimeoutErrorlong reasoning (32,516 chars) ×4write-guard blocked a truncated write ×2547
🔍 view

torch-tensor-parallelism — 0/1 passed

💬 analyst comment

New recovery blind spot: an 89,121-char visible-TEXT reasoning block in a turn that still ended in a toolCall (stop=toolUse). The runaway recovery only fires on no-toolCall turns, so this replayed ~22k tokens into every later request. Trial also went to lg:11 and was ended by the loop-guard hard-stop.

#resulttotalagentin/out tokflags
1FAIL34m33s28m25s7519104/182399
loop-guard hard-stopped the sessionempty final message (no text, no tool call)long reasoning (24,062 chars) ×7loop-guard blocked a repeated call ×11loop-guard escalation nudge
🔍 view

train-fasttext — 0/1 passed

#resulttotalagentin/out tokflags
1ERR2h00m2h00m5394213/21522
trial errored: AgentTimeoutErrora bash command timed out ×2
🔍 view

tune-mjcf — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL22m15s21m38s8208114/178681
loop-guard hard-stopped the sessionlong reasoning (32,306 chars) ×9loop-guard blocked a repeated call ×10loop-guard escalation nudgerunaway / empty-final recovery fired
🔍 view

video-processing — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL8m11s7m28s3239480/50388
long reasoning (26,169 chars) ×2
🔍 view

vulnerable-secret — 1/1 passed

#resulttotalagentin/out tokflags
1PASS3m12s2m38s317947/26563
long reasoning (15,346 chars)
🔍 view

winning-avg-corewars — 0/1 passed

💬 analyst comment

Loop-guard escalation worked end-to-end here: 38 blocks, 2 nudges, then hard-stop at 45m — one of 7 trials where the hard-stop fired (first suite-scale validation of the mechanism). Still one assertion away on the verifier.

#resulttotalagentin/out tokflags
1FAIL46m06s45m30s37788187/324880
loop-guard hard-stopped the sessionempty final message (no text, no tool call) ×2long reasoning (28,676 chars) ×13loop-guard blocked a repeated call ×38loop-guard escalation nudge ×2
🔍 view

write-compressor — 0/1 passed

#resulttotalagentin/out tokflags
1ERR30m52s30m00s597251/61558
trial errored: AgentTimeoutErrorlong reasoning (20,435 chars) ×6
🔍 view