← older: resource__qwen3.6-35b-a3b__20260721-194555all runs · task boardnewer: smoke__qwen3.6-35b-a3b__20260726-002616

suite__qwen3.6-35b-a3b__20260724-160500

Full suite — RESOURCE-LIFTED arm (24 CPU / 32 GB)

This is NOT a standard suite run. It re-runs all 89 tasks with the container resource caps lifted to CPU_OVERRIDE=24 / MEM_OVERRIDE_MB=32768, to measure the full-suite effect of fixing the environment/resource-bound tasks. Everything else matches the 40/89 baseline (suite__20260720-025830): MinimalPi, qwen3.6-35b-a3b, K=1, AGENT_TIMEOUT_MULT=2.0, all config/harness.env defaults.

NON-STANDARD / not leaderboard-comparable: Harbor's leaderboard static-validation rejects override_cpus + override_memory_mb (recorded in config.json). Compare its per-task board dots against the standard suite runs as an A/B, not as a continuation of the standard line.

Baseline to beat: 40/89 = 44.9% (suite__20260720-025830, standard caps).

Prior 7-task resource probe (runs/resource__qwen3.6-35b-a3b__20260721-194555) flipped 3 of the env-bound tasks to PASS (caffe-cifar-10, mcmc-sampling-stan, qemu-alpine-ssh) → implied ~43/89. The other 4 stayed non-pass for non-resource reasons (install-windows-3.11 GUI-capability, make-doom-for-mips hard cross-compile, train-fasttext long multi-run compute, extract-moves-from-video long+wrong-output).

RESULT: 31/89 = 34.8% (vs 40/89 baseline). Elapsed 30h33m. 17 regressions, 8 gains.

This is a LOW K=1 VARIANCE DRAW, not resource harm. Three K=1 MinimalPi/35b points now: 35 (0718) / 40 (0720) / 31 (0724) → mean ~35, spread ±~5. The 40 baseline was a high draw; 31 is a low draw (regression to the mean explains the lopsided 17-loss/8-gain vs the high baseline).

Evidence it is NOT the override: - Regressions dominated by resource-IRRELEVANT tasks failing on task logic: nginx-request-logging, vulnerable-secret, bn-fit-modify, count-dataset-tokens, pypi-server, pytorch-model-cli (no path by which 24 CPU/32 GB breaks these). - caffe-cifar-10 PASSED in the 7-task probe with the identical override but ERR'd here — same config, opposite result = variance. - Wall-clock 30h33m < 32h09m baseline → no server starvation (would be longer). - Several ERR regressions (cobol, compile-compcert, largest-eigenval, tune-mjcf) are the reasoning-THRASH spiral (unique reworded rumination, "Wait,"x154 etc.), the K=1 variance driver — not caught by DRY/loop-guard/recovery (see AGENT_TODOS).

The override DID deliver its intended gains (mcmc-sampling-stan ERR->PASS, qemu-alpine-ssh ERR->PASS) but the ~2-3 task resource signal is swamped by the ±~8 task K=1 variance. Verdict: resource lift works; measuring it needs K>=3 or the isolated probe, not a single noisy full suite.

What to check when this finishes

Run details

modelllama-local/qwen3.6-35b-a3bagentharnesses.minimal_pi:MinimalPithinkingonreasoning budgetdirect (server/none)* — see journal for the authoritative mechanismmaxTokens / contextWindow65536 / 196608agent timeout ×2.0trials89 of 89 — 31 pass · 41 fail · 17 erroredmean reward0.35tokens (job total)603,179,631 in / 8,628,686 outstarted / finished2026-07-24T16:05 / 2026-07-25T22:38wall clock30h33m

Errored trials: AgentTimeoutError (adaptive-rejection-sampler, break-filter-js-from-html, caffe-cifar-10, cobol-modernization, code-from-image, compile-compcert, feal-linear-cryptanalysis, largest-eigenval, llm-inference-batching-scheduler, make-doom-for-mips, make-mips-interpreter, mteb-retrieve, overfull-hbox, protein-assembly, train-fasttext, tune-mjcf, write-compressor)

Tasks

adaptive-rejection-sampler — 0/1 passed

#resulttotalagentin/out tokflags
1ERR31m48s30m00s17157279/103703
trial errored: AgentTimeoutErrora bash command timed out ×2
🔍 view

bn-fit-modify — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL3m03s2m17s458351/17849
🔍 view

break-filter-js-from-html — 0/1 passed

#resulttotalagentin/out tokflags
1ERR40m37s40m00s7038574/232370
trial errored: AgentTimeoutErrorempty final message (no text, no tool call)long reasoning (22,998 chars) ×7runaway / empty-final recovery fired
🔍 view

build-cython-ext — 1/1 passed

#resulttotalagentin/out tokflags
1PASS2m59s2m29s1972105/13345
🔍 view

build-pmars — 1/1 passed

#resulttotalagentin/out tokflags
1PASS2m31s1m58s337738/15594
long reasoning (12,320 chars)
🔍 view

build-pov-ray — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL16m26s15m42s3416137/97125
loop-guard hard-stopped the sessiona bash command timed outlong reasoning (29,685 chars)loop-guard blocked a repeated call ×17loop-guard escalation nudge ×2
🔍 view

caffe-cifar-10 — 0/1 passed

#resulttotalagentin/out tokflags
1ERR40m38s40m00s718528/6914
trial errored: AgentTimeoutErrora bash command timed out
🔍 view

cancel-async-tasks — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL1m15s31s55433/5075
🔍 view

chess-best-move — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL12m43s12m05s2750557/103740
long reasoning (14,439 chars) ×8
🔍 view

circuit-fibsqrt — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL42m50s41m20s12379793/327087
loop-guard hard-stopped the sessionlong reasoning (25,931 chars) ×13loop-guard blocked a repeated call ×30loop-guard escalation nudge ×2
🔍 view

cobol-modernization — 0/1 passed

#resulttotalagentin/out tokflags
1ERR30m31s30m00s366043/45998
trial errored: AgentTimeoutErrorlong reasoning (12,198 chars) ×2
🔍 view

code-from-image — 0/1 passed

#resulttotalagentin/out tokflags
1ERR40m32s40m00s4587083/265785
trial errored: AgentTimeoutErrora bash command timed outempty final message (no text, no tool call)long reasoning (26,229 chars) ×27loop-guard blocked a repeated call ×7
🔍 view

compile-compcert — 0/1 passed

#resulttotalagentin/out tokflags
1ERR1h20m1h20m57480475/118088
trial errored: AgentTimeoutErrora bash command timed out ×2
🔍 view

configure-git-webserver — 1/1 passed

#resulttotalagentin/out tokflags
1PASS1m11s30s72410/3095
🔍 view

constraints-scheduling — 1/1 passed

#resulttotalagentin/out tokflags
1PASS2m47s2m08s393383/22902
🔍 view

count-dataset-tokens — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL4m38s4m06s338118/9404
a bash command timed out
🔍 view

crack-7z-hash — 1/1 passed

#resulttotalagentin/out tokflags
1PASS1m24s48s85706/2243
🔍 view

custom-memory-heap-crash — 1/1 passed

#resulttotalagentin/out tokflags
1PASS11m29s10m47s8694891/94152
long reasoning (13,965 chars) ×5
🔍 view

db-wal-recovery — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL12m56s12m16s4673213/109031
long reasoning (23,886 chars) ×7
🔍 view

distribution-search — 1/1 passed

#resulttotalagentin/out tokflags
1PASS3m40s3m06s290800/31833
empty final message (no text, no tool call)long reasoning (12,878 chars) ×2runaway / empty-final recovery fired
🔍 view

dna-assembly — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL31m26s30m45s2829466/184387
long reasoning (21,145 chars) ×15
🔍 view

dna-insert — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL6m42s5m59s3201010/54521
claimed success but the verifier did NOT pass (heuristic)long reasoning (13,022 chars)
🔍 view

extract-elf — 1/1 passed

#resulttotalagentin/out tokflags
1PASS4m45s4m06s735202/24674
a bash command timed out
🔍 view

extract-moves-from-video — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL41m37s40m55s10567175/62685
a bash command timed out
🔍 view

feal-differential-cryptanalysis — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL41m39s40m54s4872864/330596
loop-guard hard-stopped the sessionlong reasoning (22,181 chars) ×10loop-guard blocked a repeated call ×15loop-guard escalation nudge
🔍 view

feal-linear-cryptanalysis — 0/1 passed

#resulttotalagentin/out tokflags
1ERR1h00m1h00m6594212/279932
trial errored: AgentTimeoutErrorempty final message (no text, no tool call) ×2long reasoning (24,049 chars) ×15runaway / empty-final recovery fired ×2write-guard blocked a truncated write
🔍 view

filter-js-from-html — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL3m42s11s13661/1768
🔍 view

financial-document-processor — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL13m45s13m01s8900470/84485
claimed success but the verifier did NOT pass (heuristic)
🔍 view

fix-code-vulnerability — 1/1 passed

#resulttotalagentin/out tokflags
1PASS1m00s30s200547/3615
🔍 view

fix-git — 1/1 passed

#resulttotalagentin/out tokflags
1PASS50s19s51551/2653
🔍 view

fix-ocaml-gc — 1/1 passed

#resulttotalagentin/out tokflags
1PASS7m37s3m59s492980/19295
long reasoning (24,693 chars) ×2
🔍 view

gcode-to-text — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL30m14s29m42s4835082/220843
loop-guard hard-stopped the sessionempty final message (no text, no tool call)long reasoning (15,927 chars) ×24loop-guard blocked a repeated call ×13loop-guard escalation nudge
🔍 view

git-leak-recovery — 1/1 passed

#resulttotalagentin/out tokflags
1PASS57s14s29321/2206
🔍 view

git-multibranch — 1/1 passed

#resulttotalagentin/out tokflags
1PASS2m08s1m23s389211/11496
🔍 view

gpt2-codegolf — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL21m16s20m39s4822876/172259
long reasoning (25,193 chars) ×10loop-guard blocked a repeated call
🔍 view

headless-terminal — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL6m44s6m09s1690416/30232
a bash command timed out ×2
🔍 view

hf-model-inference — 1/1 passed

#resulttotalagentin/out tokflags
1PASS5m11s4m39s1469347/33903
long reasoning (12,409 chars)
🔍 view

install-windows-3.11 — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL54m43s53m51s11497537/69692
a bash command timed out ×3
🔍 view

kv-store-grpc — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL1m03s32s125261/4199
claimed success but the verifier did NOT pass (heuristic)
🔍 view

large-scale-text-editing — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL8m29s7m14s1590943/29444
claimed success but the verifier did NOT pass (heuristic)a bash command timed outlong reasoning (18,420 chars)
🔍 view

largest-eigenval — 0/1 passed

#resulttotalagentin/out tokflags
1ERR30m33s30m00s13970989/133117
trial errored: AgentTimeoutErrorlong reasoning (29,087 chars) ×6
🔍 view

llm-inference-batching-scheduler — 0/1 passed

#resulttotalagentin/out tokflags
1ERR1h00m1h00m62730954/398478
trial errored: AgentTimeoutErrorlong reasoning (14,392 chars) ×14loop-guard blocked a repeated call ×5
🔍 view

log-summary-date-ranges — 1/1 passed

#resulttotalagentin/out tokflags
1PASS47s15s41925/2149
🔍 view

mailman — 1/1 passed

#resulttotalagentin/out tokflags
1PASS4m04s3m12s2205280/20898
long reasoning (16,968 chars)
🔍 view

make-doom-for-mips — 0/1 passed

#resulttotalagentin/out tokflags
1ERR31m07s30m00s41303651/159565
trial errored: AgentTimeoutErrora bash command timed outlong reasoning (17,043 chars)
🔍 view

make-mips-interpreter — 0/1 passed

#resulttotalagentin/out tokflags
1ERR1h01m1h00m25215800/415781
trial errored: AgentTimeoutErrorlong reasoning (12,563 chars) ×7
🔍 view

mcmc-sampling-stan — 1/1 passed

#resulttotalagentin/out tokflags
1PASS28m48s16m44s242255/9181
🔍 view

merge-diff-arc-agi-task — 1/1 passed

#resulttotalagentin/out tokflags
1PASS1m57s1m16s291004/11896
🔍 view

model-extraction-relu-logits — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL26m38s22m16s1967678/66539
claimed success but the verifier did NOT pass (heuristic)a bash command timed outlong reasoning (30,027 chars) ×5
🔍 view

modernize-scientific-stack — 1/1 passed

#resulttotalagentin/out tokflags
1PASS55s14s19412/2115
🔍 view

mteb-leaderboard — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL9m11s8m34s6204226/20966
🔍 view

mteb-retrieve — 0/1 passed

#resulttotalagentin/out tokflags
1ERR1h04m1h00m10517027/478664
trial errored: AgentTimeoutErrorlong reasoning (14,675 chars) ×9loop-guard blocked a repeated call ×3
🔍 view

multi-source-data-merger — 1/1 passed

#resulttotalagentin/out tokflags
1PASS1m06s26s34137/4192
🔍 view

nginx-request-logging — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL58s24s75338/3399
🔍 view

openssl-selfsigned-cert — 1/1 passed

#resulttotalagentin/out tokflags
1PASS1m01s30s45806/5213
🔍 view

overfull-hbox — 0/1 passed

#resulttotalagentin/out tokflags
1ERR25m55s25m00s14108424/201807
trial errored: AgentTimeoutErrorlong reasoning (26,084 chars) ×22
🔍 view

password-recovery — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL21m51s20m11s6767599/160727
loop-guard hard-stopped the sessionlong reasoning (12,776 chars) ×5loop-guard blocked a repeated call ×10loop-guard escalation nudge
🔍 view

path-tracing — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL49m53s49m04s14058944/304821
loop-guard hard-stopped the sessiona bash command timed out ×5long reasoning (13,727 chars) ×6loop-guard blocked a repeated call ×14loop-guard escalation nudge
🔍 view

path-tracing-reverse — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL44m28s43m47s31301073/219509
long reasoning (15,478 chars) ×7loop-guard blocked a repeated call
🔍 view

polyglot-c-py — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL11m14s9m59s1307319/91979
long reasoning (25,750 chars) ×10
🔍 view

polyglot-rust-c — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL22m27s21m45s4269625/190146
loop-guard hard-stopped the sessionlong reasoning (27,465 chars) ×13loop-guard blocked a repeated call ×11loop-guard escalation nudgerunaway / empty-final recovery fired
🔍 view

portfolio-optimization — 1/1 passed

#resulttotalagentin/out tokflags
1PASS3m00s1m39s285109/8143
🔍 view

protein-assembly — 0/1 passed

#resulttotalagentin/out tokflags
1ERR1h00m1h00m6488283/90329
trial errored: AgentTimeoutErrora bash command timed out ×3long reasoning (30,967 chars) ×4runaway / empty-final recovery fired
🔍 view

prove-plus-comm — 1/1 passed

#resulttotalagentin/out tokflags
1PASS1m47s1m06s96735/11140
long reasoning (15,300 chars)
🔍 view

pypi-server — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL54s22s80940/2358
claimed success but the verifier did NOT pass (heuristic)
🔍 view

pytorch-model-cli — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL7m53s6m07s537359/11945
🔍 view

pytorch-model-recovery — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL5m56s1m04s105409/9257
claimed success but the verifier did NOT pass (heuristic)long reasoning (17,410 chars)
🔍 view

qemu-alpine-ssh — 1/1 passed

#resulttotalagentin/out tokflags
1PASS25m00s24m26s1845273/35505
a bash command timed out ×7
🔍 view

qemu-startup — 1/1 passed

#resulttotalagentin/out tokflags
1PASS1m53s1m16s84460/4415
🔍 view

query-optimize — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL15m44s8m39s258678/10542
claimed success but the verifier did NOT pass (heuristic)a bash command timed out
🔍 view

raman-fitting — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL7m43s7m12s4383716/65130
claimed success but the verifier did NOT pass (heuristic)long reasoning (14,356 chars)
🔍 view

regex-chess — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL55m45s55m12s16383421/421098
loop-guard hard-stopped the sessionlong reasoning (32,576 chars) ×25loop-guard blocked a repeated call ×21loop-guard escalation nudge ×2
🔍 view

regex-log — 1/1 passed

#resulttotalagentin/out tokflags
1PASS19m58s18m34s4824859/155431
long reasoning (16,647 chars) ×6
🔍 view

reshard-c4-data — 1/1 passed

#resulttotalagentin/out tokflags
1PASS3m46s2m17s820208/17207
🔍 view

rstan-to-pystan — 1/1 passed

#resulttotalagentin/out tokflags
1PASS16m26s15m45s1374024/15047
a bash command timed out
🔍 view

sam-cell-seg — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL12m12s10m48s3502143/43437
a bash command timed out
🔍 view

sanitize-git-repo — 1/1 passed

#resulttotalagentin/out tokflags
1PASS2m13s1m41s1193745/14175
🔍 view

schemelike-metacircular-eval — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL45m38s45m03s65193522/285063
long reasoning (32,351 chars) ×11
🔍 view

sparql-university — 1/1 passed

#resulttotalagentin/out tokflags
1PASS11m03s10m21s9924602/82834
long reasoning (17,377 chars) ×3
🔍 view

sqlite-db-truncate — 1/1 passed

#resulttotalagentin/out tokflags
1PASS10m14s9m42s1091517/87629
long reasoning (13,494 chars) ×8
🔍 view

sqlite-with-gcov — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL1m44s1m01s69808/2510
🔍 view

torch-pipeline-parallelism — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL29m33s24m14s13935711/113600
a bash command timed outlong reasoning (33,389 chars) ×5loop-guard blocked a repeated call
🔍 view

torch-tensor-parallelism — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL22m11s16m20s2418977/97314
a bash command timed outlong reasoning (33,132 chars) ×5
🔍 view

train-fasttext — 0/1 passed

#resulttotalagentin/out tokflags
1ERR2h00m2h00m4173131/38441
trial errored: AgentTimeoutErrora bash command timed out ×4empty final message (no text, no tool call)runaway / empty-final recovery fired
🔍 view

tune-mjcf — 0/1 passed

#resulttotalagentin/out tokflags
1ERR30m46s30m00s6726612/176986
trial errored: AgentTimeoutErrorlong reasoning (19,941 chars) ×16
🔍 view

video-processing — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL6m31s5m49s1942606/53169
long reasoning (12,468 chars) ×3
🔍 view

vulnerable-secret — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL13m40s13m05s1698138/115090
claimed success but the verifier did NOT pass (heuristic)long reasoning (17,669 chars) ×11
🔍 view

winning-avg-corewars — 0/1 passed

#resulttotalagentin/out tokflags
1FAIL40m39s40m05s19967538/299884
loop-guard hard-stopped the sessionlong reasoning (25,705 chars) ×23loop-guard blocked a repeated call ×11loop-guard escalation nudge
🔍 view

write-compressor — 0/1 passed

#resulttotalagentin/out tokflags
1ERR30m40s30m00s8918892/211647
trial errored: AgentTimeoutErrorempty final message (no text, no tool call)long reasoning (25,219 chars) ×14runaway / empty-final recovery fired
🔍 view