terminal-bench on RTX 3090 — task status

📋 62 open agent todos (56 done) — follow-ups & proposed changes from run analyses, from AGENT_TODOS.md.

📊 Analysis & experiment write-ups: MinimalPi harness tweaks — inventory + usefulness (2026-07-21) · Full-suite analysis — qwen3.6-35b-a3b MoE on terminal-bench 2.0 (89/89 complete) · Failure-mode analysis — qwen3.6-35b-a3b MoE on terminal-bench 2.0 (partial suite) · Experiment 2026-07-05 — write guard + min-p 0.0: combined smoke measurement · Experiment 2026-07-06 — first `make fast` run: fast-fail timeouts + the 5 new quick-fail tasks · MTP speculative decoding vs context depth (qwen3.6-27b / 35b-a3b)

Status of every task in terminal-bench 2.0 when run by local Qwen models (llama.cpp on a single RTX 3090). Pick a model and harness config to keep fast and full-suite results from the same board; missing rows mean that combination has not run that task yet. Results are compiled across every finished run of that model+harness config: the badge is the latest run's verdict, the dots next to it are the per-run history (oldest → newest, one dot per run, click a dot for that run's page) with a passed/total count — so a task that passed before but fails now shows both. The task sets on top are the iteration loop (make fast, make run SET=fast2, …: quick task subsets with fast-fail timeouts, defined in config/task_sets.env); below them, the rest of the suite compiled from the full-suite runs. Click a column header to sort; the set chips filter to tasks covered by one set's runs; mixed history filters to tasks whose runs disagree; analyzed links to the per-task failure analysis. Run history & the tuning story live on the runs page.

model:harness:set:result:analysis:

Task sets — the iteration loop

Qwen 3.6 35B A3B · pi · status compiled from all 3 fast runs (newest fast__qwen3.6-35b-a3b__20260710-213727) · all 3 fast2 runs (newest fast2__qwen3.6-35b-a3b__20260710-192053) · baseline from all 4 suite runs (newest suite__qwen3.6-35b-a3b__20260724-160500) · showing 15 of 15

tasksetset runssuite baselineagentanalysis
cancel-async-tasksfastFAIL cut@3m00s 0/3FAIL 0/43m00sanalyzed ✓
filter-js-from-htmlfastFAIL 0/3FAIL 0/440sanalyzed ✓
financial-document-processorfast2FAIL 1/3FAIL 1/414m49snot analyzed
git-multibranchfast2FAIL 0/3PASS 1/42m02sanalyzed ✓
headless-terminalfast2FAIL 2/3FAIL 2/43m42snot analyzed
mailmanfast2FAIL 1/3PASS 2/44m47snot analyzed
query-optimizefastFAIL 0/3FAIL 1/45m05sanalyzed ✓
regex-logfastFAIL cut@10m00s 2/3PASS 4/410m00snot analyzed
reshard-c4-datafast2FAIL cut@5m00s 1/3PASS 3/46m17snot analyzed
fix-gitfastPASS 3/3PASS 3/48snot analyzed
fix-ocaml-gcfast2PASS 1/3PASS 3/47m43snot analyzed
nginx-request-loggingfastPASS 3/3FAIL 3/424snot analyzed
openssl-selfsigned-certfastPASS 3/3PASS 3/415snot analyzed
sanitize-git-repofastPASS cut@3m00s 3/3PASS 2/43m02snot analyzed
sparql-universityfastPASS 1/3PASS 1/41m17sanalyzed ✓

Rest of the suite

Qwen 3.6 35B A3B · pi · status compiled from all 4 suite runs (newest suite__qwen3.6-35b-a3b__20260724-160500) · showing 74 of 74

tasksuite runsagentanalysis
bn-fit-modifyFAIL 1/42m17snot analyzed
build-pov-rayFAIL 2/415m42snot analyzed
chess-best-moveFAIL 0/412m05snot analyzed
circuit-fibsqrtFAIL 0/441m20snot analyzed
count-dataset-tokensFAIL 2/44m06snot analyzed
db-wal-recoveryFAIL 0/412m16snot analyzed
dna-assemblyFAIL 0/430m45snot analyzed
dna-insertFAIL 0/45m59snot analyzed
extract-moves-from-videoFAIL 0/440m55snot analyzed
feal-differential-cryptanalysisFAIL 1/440m54snot analyzed
gcode-to-textFAIL 0/429m42snot analyzed
gpt2-codegolfFAIL 0/420m39snot analyzed
install-windows-3.11FAIL 0/453m51snot analyzed
kv-store-grpcFAIL 2/432snot analyzed
large-scale-text-editingFAIL 2/47m14snot analyzed
model-extraction-relu-logitsFAIL 0/422m16snot analyzed
mteb-leaderboardFAIL 0/48m34snot analyzed
password-recoveryFAIL 0/420m11snot analyzed
path-tracingFAIL 0/449m04snot analyzed
path-tracing-reverseFAIL 0/443m47snot analyzed
polyglot-c-pyFAIL 0/49m59snot analyzed
polyglot-rust-cFAIL 0/421m45snot analyzed
pypi-serverFAIL 3/422snot analyzed
pytorch-model-cliFAIL 3/46m07snot analyzed
pytorch-model-recoveryFAIL 1/41m04snot analyzed
raman-fittingFAIL 0/47m12snot analyzed
regex-chessFAIL 0/455m12snot analyzed
sam-cell-segFAIL 0/410m48snot analyzed
schemelike-metacircular-evalFAIL 0/445m03snot analyzed
sqlite-with-gcovFAIL 1/41m01snot analyzed
torch-pipeline-parallelismFAIL 0/424m14snot analyzed
torch-tensor-parallelismFAIL 1/416m20snot analyzed
video-processingFAIL 0/45m49snot analyzed
vulnerable-secretFAIL 3/413m05snot analyzed
winning-avg-corewarsFAIL 0/440m05snot analyzed
adaptive-rejection-samplerERR 0/430m00snot analyzed
break-filter-js-from-htmlERR 1/440m00snot analyzed
caffe-cifar-10ERR 0/440m00snot analyzed
cobol-modernizationERR 3/430m00snot analyzed
code-from-imageERR 2/440m00snot analyzed
compile-compcertERR 3/41h20mnot analyzed
feal-linear-cryptanalysisERR 1/41h00mnot analyzed
largest-eigenvalERR 2/430m00snot analyzed
llm-inference-batching-schedulerERR 1/41h00mnot analyzed
make-doom-for-mipsERR 0/430m00snot analyzed
make-mips-interpreterERR 0/41h00mnot analyzed
mteb-retrieveERR 0/41h00mnot analyzed
overfull-hboxERR 1/425m00snot analyzed
protein-assemblyERR 0/41h00mnot analyzed
train-fasttextERR 0/42h00mnot analyzed
tune-mjcfERR 2/430m00snot analyzed
write-compressorERR 0/430m00snot analyzed
build-cython-extPASS 1/42m29snot analyzed
build-pmarsPASS 2/41m58snot analyzed
configure-git-webserverPASS 4/430snot analyzed
constraints-schedulingPASS 4/42m08snot analyzed
crack-7z-hashPASS 2/448snot analyzed
custom-memory-heap-crashPASS 4/410m47snot analyzed
distribution-searchPASS 4/43m06snot analyzed
extract-elfPASS 4/44m06snot analyzed
fix-code-vulnerabilityPASS 4/430snot analyzed
git-leak-recoveryPASS 4/414snot analyzed
hf-model-inferencePASS 4/44m39snot analyzed
log-summary-date-rangesPASS 4/415snot analyzed
mcmc-sampling-stanPASS 3/416m44snot analyzed
merge-diff-arc-agi-taskPASS 3/41m16snot analyzed
modernize-scientific-stackPASS 4/414snot analyzed
multi-source-data-mergerPASS 4/426snot analyzed
portfolio-optimizationPASS 4/41m39snot analyzed
prove-plus-commPASS 4/41m06snot analyzed
qemu-alpine-sshPASS 1/424m26snot analyzed
qemu-startupPASS 2/41m16snot analyzed
rstan-to-pystanPASS 3/415m45snot analyzed
sqlite-db-truncatePASS 1/49m42snot analyzed

15 in task sets · 74 suite tasks · 5 analyzed · 52 with mixed history · analyses live in analysis/qwen3.6-35b-a3b/pi/<task>.md