Overview 152 result records in results, latest run 2026-10-05 00:18 AEST. Every number on this site comes from those records. Failed and crashed runs are listed alongside the passing ones.
Qwen3.8-Flash-Next running 59 records, no qualified eligible config yet
Target: lab decode ≥ 50 tok/s and lab prefill > 1000 tok/s on the same config, passing §5 stability; kit band top-1 ≥ 0.987, KL ≤ 0.0013.
No qualified headline yet. Best so far
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 · current · x8 · EXL3 3.05 bpw
59 runs in 14 configs: ✓ pass 45 ! fail 8 ✕ crash 6 · leaderboard
89 records, no qualified eligible config yet
Target: 1× RTX 3090, registry speed gate 15 tok/s (lab C1 decode), §5 stability and quality floor (≥ 80% of the reference solve rate). 3-card configs are a separate recipe.
Best 1 card
Config ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 · ik_llama.cpp · GGUF IQ2_M · current · x8 · below 15 tok/s: documented experiment, no registry PR
Best 3 cards
Config ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 · ik_llama.cpp · GGUF IQ2_M · current · 3× x8 · below 15 tok/s: documented experiment, no registry PR
89 runs in 22 configs: ✓ pass 39 ! fail 43 ✕ crash 7 · leaderboard
Pull requests PRs open in Phase 3, after the operator review of this site.
Recent runs 64 failed or crashed runs: see Runs and dropped configs .
Qwen3.8-Flash-Next leaderboard One row per config_id, aggregating all of its records. Decode/prefill are the best value from a passing record (from a failed one when none passed, marked † and never counted as meeting a target). ✓ = meets the target (decode ≥ 50 tok/s, prefill > 1000 tok/s). Chips link to each run; ! = fail, ✕ = crash, – = skipped.
Config Lab decode tok/s Lab prefill tok/s TTFT prefill tok/s Kit C1 tok/s Occupied max Ctx configured Quality Stability Tool calls Engine Quant Cards Layout · PCIe Eligibility Runs sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 ✕ 1 crash ! 1 fail 57.4 ✓ 2,537 ✓ 2,009 47.4 167k 205k in band top-1 0.98967 · KL 0.00097327 failed 2 runs · min win 39.7 71 ✕ sglang EXL3 3.05 bpw 1 current · x8 none load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012 ! 1 fail 57.1 ✓ 608 ✕ 987 – 106k 131k – failed 1 run · min win 38.6 70 ✕ llama.cpp GGUF IQ4_XS 3 current · 3× x8 none load lab lab TTFT stability ! sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 ! 1 fail 55.6 ✓ 2,596 ✓ 2,017 48.2 167k 205k in band top-1 0.98867 · KL 0.00096647 failed 1 run · min win 37.9 67 ✕ sglang EXL3 3.05 bpw 1 current · x8 none load lab lab kit sweep TTFT stability ! quality panel sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2 ! 1 fail 55.6 ✓ 2,580 ✓ 2,014 – 167k 205k – failed 1 run · min win 40.1 57 ✕ sglang EXL3 3.05 bpw 1 current · x8 none load lab lab TTFT stability ! fn-trellis-3.05-0xsero-main-x8-gpu2 ! 1 fail 55.3 ✓ 2,228 ✓ 2,263 46.9 167k 205k in band top-1 0.99234 · KL 0.0009667 failed 1 run · min win 38.1 63 ✕ sglang EXL3 3.05 bpw 1 current · x8 none load lab lab kit sweep TTFT stability ! quality panel sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2 ! 1 fail 53.7 ✓ 2,596 ✓ 2,022 – 167k 205k – failed 1 run · min win 40.9 72 ✕ sglang EXL3 3.05 bpw 1 current · x8 none load lab lab TTFT stability ! llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012 44.3 ✕ 717 ✕ 1,067 – 106k 131k – – – llama.cpp GGUF IQ4_XS 3 current · 3× x8 none load lab TTFT llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012 35.0 ✕ 713 ✕ 980 – 106k 131k – – – llama.cpp GGUF IQ4_XS 3 current · 3× x8 none load lab TTFT fn-llamacpp-iq4xs-128k-1x-fit-gpu2 ! 1 fail 26.5 ✕ 157 ✕ 269 – 106k 131k – failed 1 run · min win 23.8 49 ✕ llama.cpp GGUF IQ4_XS 1 current · x8 none load lab lab TTFT stability ! fn-llamacpp-iq4xs-128k-1x-gpu2 ✕ 1 crash – – – – – 131k – – – llama.cpp GGUF IQ4_XS 1 current · x8 none load ✕ llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012 ! 1 fail – – – – – 131k – failed 1 run 0 ✕ llama.cpp GGUF IQ4_XS 3 current · 3× x8 none load stability ! sglang-qwen3.8-flash-next-exl3-2.05bpw-offload-200k-rtx-3090-24gb-gpu2 ✕ 1 crash – – – – – 205k – – – sglang EXL3 2.05 bpw 1 current · x8 none load ✕ sglang-qwen3.8-flash-next-exl3-4.05bpw-offload-200k-rtx-3090-24gb-gpu2 ✕ 1 crash – – – – – 205k – – – sglang EXL3 4.05 bpw 1 current · x8 none load ✕ sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2 ✕ 2 crash – – – – – 205k – – – sglang EXL3 3.05 bpw 1 current · x8 none load lab ✕ TTFT ✕
GLM-5.3-Flash leaderboard One row per config_id, aggregating all of its records. Decode/prefill are the best value from a passing record (from a failed one when none passed, marked † and never counted as meeting a target). ✓ = meets the target (decode ≥ 15 tok/s). Chips link to each run; ! = fail, ✕ = crash, – = skipped.
Charts Each chart appears once records contain its measurements. Hover or tap a mark for its run; every chart has a table view.
Decode vs occupied context Lab C1 decode and the stability minimum 60 s window against the occupied prompt tokens reported by the server.
Qwen3.8-Flash-Next GLM-5.3-Flash lab C1 (first 30 s) stability: min 60 s window0 20 40 60 80 0 50 100 150 200 occupied context (k tokens) decode tok/s Flash-Next target 50 GLM gate 15 128k fn-trellis-3.05-0xsero-main-x8-gpu2 · lab C1 decode 49.3 tok/s at 166,667 occupied · pass fn-trellis-3.05-0xsero-main-x8-gpu2 · stability min 60 s window 38.1 tok/s at 30,034 occupied · fail fn-trellis-3.05-0xsero-main-x8-gpu2 · lab C1 decode 55.3 tok/s at 166,667 occupied · pass fn-llamacpp-iq4xs-128k-1x-fit-gpu2 · lab C1 decode 26.5 tok/s at 106,294 occupied · pass fn-llamacpp-iq4xs-128k-1x-fit-gpu2 · stability min 60 s window 23.8 tok/s at 16,819 occupied · fail fn-llamacpp-iq4xs-128k-1x-fit-gpu2 · lab C1 decode 26.1 tok/s at 106,294 occupied · pass glm-llamacpp-iq2m-128k-1x-gpu2 · lab C1 decode 9.60 tok/s at 93,437 occupied · fail glm-llamacpp-iq2m-128k-1x-gpu2 · stability min 60 s window 8.91 tok/s at 9,047 occupied · fail glm-llamacpp-iq2m-128k-1x-gpu2 · lab C1 decode 9.60 tok/s at 93,437 occupied · fail llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 · lab C1 decode 9.20 tok/s at 93,437 occupied · fail llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 · stability min 60 s window 8.80 tok/s at 6,658 occupied · fail llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 · lab C1 decode 9.30 tok/s at 93,437 occupied · fail ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 · lab C1 decode 10.0 tok/s at 93,437 occupied · fail ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 · stability min 60 s window 9.46 tok/s at 1,874 occupied · fail ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 · lab C1 decode 9.50 tok/s at 93,437 occupied · fail ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2 · lab C1 decode 8.20 tok/s at 93,437 occupied · fail llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 · lab C1 decode 8.30 tok/s at 93,437 occupied · fail llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 · stability min 60 s window 7.65 tok/s at 4,875 occupied · fail llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 · lab C1 decode 8.40 tok/s at 93,437 occupied · fail llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 · lab C1 decode 9.50 tok/s at 93,437 occupied · fail llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 · stability min 60 s window 8.85 tok/s at 9,732 occupied · fail llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 · lab C1 decode 9.50 tok/s at 93,437 occupied · fail ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2 · lab C1 decode 8.60 tok/s at 93,437 occupied · fail llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 · lab C1 decode 12.3 tok/s at 93,437 occupied · fail llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 · stability min 60 s window 11.4 tok/s at 8,118 occupied · fail llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 · lab C1 decode 10.5 tok/s at 93,437 occupied · fail llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 · stability min 60 s window 9.86 tok/s at 5,309 occupied · fail llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 · lab C1 decode 10.4 tok/s at 93,437 occupied · fail llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012 · lab C1 decode 12.2 tok/s at 93,437 occupied · fail llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012 · lab C1 decode 11.8 tok/s at 93,437 occupied · fail llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012 · stability min 60 s window 11.0 tok/s at 6,174 occupied · fail llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012 · lab C1 decode 11.8 tok/s at 93,437 occupied · fail llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2 · lab C1 decode 9.50 tok/s at 93,437 occupied · fail llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 · lab C1 decode 9.60 tok/s at 93,437 occupied · fail llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 · stability min 60 s window 8.85 tok/s at 6,632 occupied · fail llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 · lab C1 decode 9.60 tok/s at 93,437 occupied · fail llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012 · lab C1 decode 8.00 tok/s at 93,437 occupied · fail llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012 · stability min 60 s window 7.81 tok/s at 6,852 occupied · fail llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012 · lab C1 decode 8.30 tok/s at 93,437 occupied · fail llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012 · lab C1 decode 56.9 tok/s at 106,294 occupied · pass llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012 · stability min 60 s window 38.6 tok/s at 39,226 occupied · fail llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012 · lab C1 decode 57.1 tok/s at 106,294 occupied · pass llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2 · lab C1 decode 9.10 tok/s at 93,437 occupied · fail llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012 · lab C1 decode 11.2 tok/s at 93,437 occupied · fail ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 · lab C1 decode 14.0 tok/s at 93,437 occupied · fail ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 · stability min 60 s window 13.3 tok/s at 1,874 occupied · fail ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 · lab C1 decode 13.4 tok/s at 93,437 occupied · fail llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012 · lab C1 decode 44.3 tok/s at 106,294 occupied · pass llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012 · lab C1 decode 35.0 tok/s at 106,294 occupied · pass ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012 · lab C1 decode 10.2 tok/s at 93,437 occupied · fail sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 · lab C1 decode 55.6 tok/s at 166,667 occupied · pass sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 · stability min 60 s window 37.9 tok/s at 36,327 occupied · fail sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 · lab C1 decode 54.5 tok/s at 166,667 occupied · pass exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 · lab C1 decode 7.40 tok/s at 93,437 occupied · fail exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 · stability min 60 s window 3.36 tok/s at 4,589 occupied · fail exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 · lab C1 decode 7.50 tok/s at 93,437 occupied · fail sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 · lab C1 decode 55.8 tok/s at 166,667 occupied · pass sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 · stability min 60 s window 39.7 tok/s at 25,309 occupied · fail sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 · lab C1 decode 57.4 tok/s at 166,667 occupied · pass sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2 · lab C1 decode 55.6 tok/s at 166,667 occupied · pass sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2 · stability min 60 s window 40.1 tok/s at 29,027 occupied · fail sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2 · lab C1 decode 54.6 tok/s at 166,667 occupied · pass sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2 · lab C1 decode 53.7 tok/s at 166,667 occupied · pass sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2 · stability min 60 s window 40.9 tok/s at 39,281 occupied · fail sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2 · lab C1 decode 51.3 tok/s at 166,667 occupied · pass Table view Model Config Measure Occupied tok/s Outcome Qwen3.8-Flash-Next fn-trellis-3.05-0xsero-main-x8-gpu2 lab C1 decode 166,667 49.3 ✓ pass Qwen3.8-Flash-Next fn-trellis-3.05-0xsero-main-x8-gpu2 stability min 60 s window 30,034 38.1 ! fail Qwen3.8-Flash-Next fn-trellis-3.05-0xsero-main-x8-gpu2 lab C1 decode 166,667 55.3 ✓ pass Qwen3.8-Flash-Next fn-llamacpp-iq4xs-128k-1x-fit-gpu2 lab C1 decode 106,294 26.5 ✓ pass Qwen3.8-Flash-Next fn-llamacpp-iq4xs-128k-1x-fit-gpu2 stability min 60 s window 16,819 23.8 ! fail Qwen3.8-Flash-Next fn-llamacpp-iq4xs-128k-1x-fit-gpu2 lab C1 decode 106,294 26.1 ✓ pass GLM-5.3-Flash glm-llamacpp-iq2m-128k-1x-gpu2 lab C1 decode 93,437 9.60 ! fail GLM-5.3-Flash glm-llamacpp-iq2m-128k-1x-gpu2 stability min 60 s window 9,047 8.91 ! fail GLM-5.3-Flash glm-llamacpp-iq2m-128k-1x-gpu2 lab C1 decode 93,437 9.60 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 lab C1 decode 93,437 9.20 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 stability min 60 s window 6,658 8.80 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 lab C1 decode 93,437 9.30 ! fail GLM-5.3-Flash ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 lab C1 decode 93,437 10.0 ! fail GLM-5.3-Flash ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 stability min 60 s window 1,874 9.46 ! fail GLM-5.3-Flash ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 lab C1 decode 93,437 9.50 ! fail GLM-5.3-Flash ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2 lab C1 decode 93,437 8.20 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 lab C1 decode 93,437 8.30 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 stability min 60 s window 4,875 7.65 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 lab C1 decode 93,437 8.40 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 lab C1 decode 93,437 9.50 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 stability min 60 s window 9,732 8.85 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 lab C1 decode 93,437 9.50 ! fail GLM-5.3-Flash ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2 lab C1 decode 93,437 8.60 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 lab C1 decode 93,437 12.3 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 stability min 60 s window 8,118 11.4 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 lab C1 decode 93,437 10.5 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 stability min 60 s window 5,309 9.86 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 lab C1 decode 93,437 10.4 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012 lab C1 decode 93,437 12.2 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012 lab C1 decode 93,437 11.8 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012 stability min 60 s window 6,174 11.0 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012 lab C1 decode 93,437 11.8 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2 lab C1 decode 93,437 9.50 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 lab C1 decode 93,437 9.60 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 stability min 60 s window 6,632 8.85 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 lab C1 decode 93,437 9.60 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012 lab C1 decode 93,437 8.00 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012 stability min 60 s window 6,852 7.81 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012 lab C1 decode 93,437 8.30 ! fail Qwen3.8-Flash-Next llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012 lab C1 decode 106,294 56.9 ✓ pass Qwen3.8-Flash-Next llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012 stability min 60 s window 39,226 38.6 ! fail Qwen3.8-Flash-Next llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012 lab C1 decode 106,294 57.1 ✓ pass GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2 lab C1 decode 93,437 9.10 ! fail GLM-5.3-Flash llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012 lab C1 decode 93,437 11.2 ! fail GLM-5.3-Flash ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 lab C1 decode 93,437 14.0 ! fail GLM-5.3-Flash ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 stability min 60 s window 1,874 13.3 ! fail GLM-5.3-Flash ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 lab C1 decode 93,437 13.4 ! fail Qwen3.8-Flash-Next llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012 lab C1 decode 106,294 44.3 ✓ pass Qwen3.8-Flash-Next llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012 lab C1 decode 106,294 35.0 ✓ pass GLM-5.3-Flash ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012 lab C1 decode 93,437 10.2 ! fail Qwen3.8-Flash-Next sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 lab C1 decode 166,667 55.6 ✓ pass Qwen3.8-Flash-Next sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 stability min 60 s window 36,327 37.9 ! fail Qwen3.8-Flash-Next sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 lab C1 decode 166,667 54.5 ✓ pass GLM-5.3-Flash exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 lab C1 decode 93,437 7.40 ! fail GLM-5.3-Flash exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 stability min 60 s window 4,589 3.36 ! fail GLM-5.3-Flash exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 lab C1 decode 93,437 7.50 ! fail Qwen3.8-Flash-Next sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 lab C1 decode 166,667 55.8 ✓ pass Qwen3.8-Flash-Next sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 stability min 60 s window 25,309 39.7 ! fail Qwen3.8-Flash-Next sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 lab C1 decode 166,667 57.4 ✓ pass Qwen3.8-Flash-Next sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2 lab C1 decode 166,667 55.6 ✓ pass Qwen3.8-Flash-Next sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2 stability min 60 s window 29,027 40.1 ! fail Qwen3.8-Flash-Next sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2 lab C1 decode 166,667 54.6 ✓ pass Qwen3.8-Flash-Next sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2 lab C1 decode 166,667 53.7 ✓ pass Qwen3.8-Flash-Next sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2 stability min 60 s window 39,281 40.9 ! fail Qwen3.8-Flash-Next sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2 lab C1 decode 166,667 51.3 ✓ pass
Speed vs quality Best lab decode per config against its quality measurement.
Qwen3.8-Flash-Next: decode vs mean KL (lower KL is better) 0 20 40 60 80 0 0.0005 0.001 0.0015 mean KL vs exllamav3 reference panel lab decode tok/s target 50 band KL ≤ 0.0013 fn-trellis-3.05-0xsero-main-x8-gpu2 · KL 0.0009667 · top-1 0.99234 · decode 55.3 fn-trellis-3.05-0xsero-main-x8-gpu2 sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 · KL 0.00096647 · top-1 0.98867 · decode 55.6 sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 · KL 0.00097327 · top-1 0.98967 · decode 57.4 sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
Table view Config KL top-1 Lab decode fn-trellis-3.05-0xsero-main-x8-gpu2 0.0009667 0.99234 55.3 sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 0.00096647 0.98867 55.6 sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 0.00097327 0.98967 57.4
1 vs 3 cards (current layout) Best config per card count, same model.
Best lab decode 0 20 40 60 80 Flash-Next EXL3 3.05 bpw · 1 card sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2: 57.4 tok/s decode 57.4 Flash-Next GGUF IQ4_XS · 3 cards llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012: 57.1 tok/s decode 57.1 GLM GGUF IQ2_M · 1 card ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2: 10.0 tok/s decode 10.0 GLM GGUF IQ2_M · 3 cards ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012: 14.0 tok/s decode 14.0 tok/s Lab prefill of the same configs 0 1k 2k 3k Flash-Next EXL3 3.05 bpw · 1 card sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2: 2,537 tok/s prefill 2,537 Flash-Next GGUF IQ4_XS · 3 cards llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012: 608 tok/s prefill 608 GLM GGUF IQ2_M · 1 card ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2: 151 tok/s prefill 151 GLM GGUF IQ2_M · 3 cards ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012: 189 tok/s prefill 189 tok/s Table view x8 vs x16 No measurements yet: needs single-card lab results in both layouts for a model.
All runs Every record, newest first, failed, crashed and skipped runs included.
Started Outcome Model Kind Config Key metrics Stage Ticket 2026-10-05 00:18 AEST ✓ pass Flash-Next lab sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 lab_decode_c1_tps 55.8 tok/s · lab_prefill_tps 2,419 tok/s T13 2026-10-05 00:14 AEST ✓ pass Flash-Next load sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 load_seconds 257 s T13 2026-10-05 00:07 AEST ✓ pass Flash-Next lab sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2 lab_decode_c1_tps 51.3 tok/s · lab_prefill_tps 2,596 tok/s T14b 2026-10-04 23:57 AEST ! fail Flash-Next stability sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2 decode_window_min_tps 40.9 tok/s · tool_calls 72 calls T14b 2026-10-04 23:55 AEST ✓ pass Flash-Next TTFT sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2 ttft_prefill_tps 2,022 tok/s T14b 2026-10-04 23:52 AEST ✓ pass Flash-Next lab sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2 lab_decode_c1_tps 53.7 tok/s · lab_prefill_tps 2,440 tok/s T14b 2026-10-04 23:48 AEST ✓ pass Flash-Next load sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2 load_seconds 215 s T14b 2026-10-04 23:43 AEST ✓ pass Flash-Next lab sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2 lab_decode_c1_tps 54.6 tok/s · lab_prefill_tps 2,580 tok/s T14a 2026-10-04 23:33 AEST ! fail Flash-Next stability sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2 decode_window_min_tps 40.1 tok/s · tool_calls 57 calls T14a 2026-10-04 23:31 AEST ✓ pass Flash-Next TTFT sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2 ttft_prefill_tps 2,014 tok/s T14a 2026-10-04 23:26 AEST ✓ pass Flash-Next lab sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2 lab_decode_c1_tps 55.6 tok/s · lab_prefill_tps 2,430 tok/s T14a 2026-10-04 23:22 AEST ✓ pass Flash-Next load sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2 load_seconds 254 s T14a 2026-10-04 23:13 AEST ✓ pass Flash-Next lab sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 lab_decode_c1_tps 57.4 tok/s · lab_prefill_tps 2,537 tok/s T14a 2026-10-04 23:02 AEST ! fail Flash-Next stability sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 decode_window_min_tps 39.7 tok/s · tool_calls 71 calls T14a 2026-10-04 22:58 AEST ✓ pass Flash-Next load sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 load_seconds 264 s T14a 2026-10-04 22:54 AEST ✓ pass GLM TTFT exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 ttft_prefill_tps 483 tok/s T23 2026-10-04 22:49 AEST ✓ pass GLM load exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 load_seconds 276 s T23 2026-10-04 22:48 AEST ✕ crash Flash-Next TTFT sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2 ttft_prefill_tps unavailable measure T14a 2026-10-04 22:47 AEST ✕ crash Flash-Next lab sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2 lab_decode_c1_tps unavailable · lab_prefill_tps unavailable harness T14a 2026-10-04 22:42 AEST ✓ pass Flash-Next load sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2 load_seconds 257 s T14a 2026-10-04 22:32 AEST ✕ crash Flash-Next stability sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 decode_window_min_tps unavailable · tool_calls 0 calls measure T13 2026-10-04 22:32 AEST ✓ pass Flash-Next quality panel sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 top1_agreement 0.98967 fraction · mean_kl 0.00097327 nats T13 2026-10-04 22:30 AEST ✓ pass Flash-Next TTFT sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 ttft_prefill_tps 2,009 tok/s T13 2026-10-04 22:15 AEST ✓ pass Flash-Next kit sweep sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 decode_c1_tps 47.4 tok/s · prefill_32k_tps 2,678 tok/s T13 2026-10-04 22:10 AEST ✓ pass Flash-Next lab sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 lab_decode_c1_tps 55.8 tok/s · lab_prefill_tps 2,419 tok/s T13 2026-10-04 22:05 AEST ✓ pass Flash-Next load sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 load_seconds 260 s T13 2026-10-04 21:27 AEST ! fail GLM lab exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 lab_decode_c1_tps 7.50 tok/s · lab_prefill_tps 720 tok/s gates T23 2026-10-04 21:17 AEST ! fail GLM stability exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 decode_window_min_tps 3.36 tok/s · tool_calls 21 calls T23 2026-10-04 21:13 AEST ! fail GLM TTFT exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 ttft_prefill_tps unavailable invalid measurement T23 2026-10-04 21:02 AEST ! fail GLM lab exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 lab_decode_c1_tps 7.40 tok/s · lab_prefill_tps 739 tok/s gates T23 2026-10-04 20:57 AEST ✓ pass GLM load exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 load_seconds 300 s T23 2026-10-04 20:56 AEST ✕ crash Flash-Next load sglang-qwen3.8-flash-next-exl3-4.05bpw-offload-200k-rtx-3090-24gb-gpu2 load_seconds unavailable load T15 2026-10-04 20:52 AEST ✕ crash Flash-Next load sglang-qwen3.8-flash-next-exl3-2.05bpw-offload-200k-rtx-3090-24gb-gpu2 load_seconds unavailable load T15 2026-10-04 20:47 AEST ✓ pass Flash-Next lab sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 lab_decode_c1_tps 54.5 tok/s · lab_prefill_tps 2,596 tok/s T13 2026-10-04 20:36 AEST ! fail Flash-Next stability sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 decode_window_min_tps 37.9 tok/s · tool_calls 67 calls T13 2026-10-04 20:36 AEST ✓ pass Flash-Next quality panel sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 top1_agreement 0.98867 fraction · mean_kl 0.00096647 nats T13 2026-10-04 20:34 AEST ✓ pass Flash-Next TTFT sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 ttft_prefill_tps 2,017 tok/s T13 2026-10-04 20:19 AEST ✓ pass Flash-Next kit sweep sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 decode_c1_tps 48.2 tok/s · prefill_32k_tps 2,678 tok/s T13 2026-10-04 20:14 AEST ✓ pass Flash-Next lab sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 lab_decode_c1_tps 55.6 tok/s · lab_prefill_tps 2,433 tok/s T13 2026-10-04 20:10 AEST ✓ pass Flash-Next load sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 load_seconds 236 s T13 2026-10-04 19:38 AEST ✕ crash GLM stability ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012 decode_window_min_tps unavailable · tool_calls 0 calls measure T22c 2026-10-04 19:26 AEST ✓ pass GLM TTFT ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012 ttft_prefill_tps 248 tok/s T22c 2026-10-04 19:12 AEST ! fail GLM lab ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012 lab_decode_c1_tps 10.2 tok/s · lab_prefill_tps 166 tok/s gates T22c 2026-10-04 19:11 AEST ✓ pass GLM load ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012 load_seconds 48.6 s T22c 2026-10-04 19:11 AEST ✓ pass none hw bw hw-bw-gpu1-current h2d_gbps 13.5 GB/s · d2h_gbps 13.2 GB/s T05a 2026-10-04 18:19 AEST ✓ pass Flash-Next TTFT llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012 ttft_prefill_tps 980 tok/s T11 2026-10-04 18:11 AEST ✓ pass Flash-Next lab llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012 lab_decode_c1_tps 35.0 tok/s · lab_prefill_tps 713 tok/s T11 2026-10-04 18:10 AEST ✓ pass Flash-Next load llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012 load_seconds 39.9 s T11 2026-10-04 18:06 AEST ✓ pass Flash-Next TTFT llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012 ttft_prefill_tps 1,067 tok/s T11 2026-10-04 18:01 AEST ✓ pass Flash-Next lab llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012 lab_decode_c1_tps 44.3 tok/s · lab_prefill_tps 717 tok/s T11 2026-10-04 18:00 AEST ✓ pass Flash-Next load llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012 load_seconds 39.9 s T11 2026-10-04 17:42 AEST ✕ crash GLM load ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm4096-gpu012 load_seconds unavailable load T22c 2026-10-04 17:31 AEST ! fail GLM lab ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 lab_decode_c1_tps 13.4 tok/s · lab_prefill_tps 189 tok/s gates T22c 2026-10-04 17:21 AEST ! fail GLM stability ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 decode_window_min_tps 13.3 tok/s · tool_calls 0 calls T22c 2026-10-04 17:10 AEST ✓ pass GLM TTFT ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 ttft_prefill_tps 283 tok/s T22c 2026-10-04 16:58 AEST ! fail GLM lab ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 lab_decode_c1_tps 14.0 tok/s · lab_prefill_tps 185 tok/s gates T22c 2026-10-04 16:58 AEST ✓ pass GLM load ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 load_seconds 44.3 s T22c 2026-10-04 16:48 AEST ✓ pass GLM TTFT llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012 ttft_prefill_tps 298 tok/s T21 2026-10-04 16:30 AEST ! fail GLM lab llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012 lab_decode_c1_tps 11.2 tok/s · lab_prefill_tps 258 tok/s gates T21 2026-10-04 16:30 AEST ✓ pass GLM load llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012 load_seconds 54.9 s T21 2026-10-04 16:23 AEST ✓ pass GLM TTFT llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2 ttft_prefill_tps 384 tok/s T20b 2026-10-04 16:14 AEST ! fail GLM lab llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2 lab_decode_c1_tps 9.10 tok/s · lab_prefill_tps 370 tok/s gates T20b 2026-10-04 16:13 AEST ✓ pass GLM load llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2 load_seconds 48.9 s T20b 2026-10-04 16:00 AEST ✓ pass Flash-Next lab llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012 lab_decode_c1_tps 57.1 tok/s · lab_prefill_tps 608 tok/s T11 2026-10-04 15:49 AEST ! fail Flash-Next stability llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012 decode_window_min_tps 38.6 tok/s · tool_calls 70 calls T11 2026-10-04 15:45 AEST ✓ pass Flash-Next TTFT llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012 ttft_prefill_tps 987 tok/s T11 2026-10-04 15:39 AEST ✓ pass Flash-Next lab llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012 lab_decode_c1_tps 56.9 tok/s · lab_prefill_tps 605 tok/s T11 2026-10-04 15:38 AEST ✓ pass Flash-Next load llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012 load_seconds 39.9 s T11 2026-10-04 15:19 AEST ! fail GLM lab llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012 lab_decode_c1_tps 8.30 tok/s · lab_prefill_tps 156 tok/s gates T21 2026-10-04 15:08 AEST ! fail GLM stability llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012 decode_window_min_tps 7.81 tok/s · tool_calls 25 calls T21 2026-10-04 14:52 AEST ✓ pass GLM TTFT llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012 ttft_prefill_tps 149 tok/s T21 2026-10-04 14:26 AEST ! fail GLM lab llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012 lab_decode_c1_tps 8.00 tok/s · lab_prefill_tps 139 tok/s gates T21 2026-10-04 14:25 AEST ✓ pass GLM load llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012 load_seconds 91.2 s T21 2026-10-04 14:24 AEST ✕ crash GLM load ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-gpu012 load_seconds unavailable load T22c 2026-10-04 14:23 AEST ✕ crash GLM load ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-gpu012 load_seconds unavailable load T22c 2026-10-04 14:14 AEST ! fail GLM lab llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 lab_decode_c1_tps 9.60 tok/s · lab_prefill_tps 288 tok/s gates T20b 2026-10-04 14:04 AEST ! fail GLM stability llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 decode_window_min_tps 8.85 tok/s · tool_calls 27 calls T20b 2026-10-04 13:55 AEST ✓ pass GLM TTFT llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 ttft_prefill_tps 321 tok/s T20b 2026-10-04 13:42 AEST ! fail GLM lab llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 lab_decode_c1_tps 9.60 tok/s · lab_prefill_tps 288 tok/s gates T20b 2026-10-04 13:41 AEST ✓ pass GLM load llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 load_seconds 45.9 s T20b 2026-10-04 13:29 AEST ✓ pass GLM TTFT llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2 ttft_prefill_tps 213 tok/s T20b 2026-10-04 13:06 AEST ! fail GLM lab llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2 lab_decode_c1_tps 9.50 tok/s · lab_prefill_tps 201 tok/s gates T20b 2026-10-04 13:05 AEST ✓ pass GLM load llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2 load_seconds 45.9 s T20b 2026-10-04 12:49 AEST ! fail GLM lab llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012 lab_decode_c1_tps 11.8 tok/s · lab_prefill_tps 218 tok/s gates T21 2026-10-04 12:39 AEST ! fail GLM stability llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012 decode_window_min_tps 11.0 tok/s · tool_calls 24 calls T21 2026-10-04 12:27 AEST ✓ pass GLM TTFT llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012 ttft_prefill_tps 240 tok/s T21 2026-10-04 12:16 AEST ! fail GLM lab llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012 lab_decode_c1_tps 11.8 tok/s · lab_prefill_tps 219 tok/s gates T21 2026-10-04 12:15 AEST ✓ pass GLM load llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012 load_seconds 54.9 s T21 2026-10-04 11:59 AEST ✓ pass GLM TTFT llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012 ttft_prefill_tps 159 tok/s T21 2026-10-04 11:39 AEST ! fail GLM lab llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012 lab_decode_c1_tps 12.2 tok/s · lab_prefill_tps 153 tok/s gates T21 2026-10-04 11:38 AEST ✓ pass GLM load llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012 load_seconds 55.0 s T21 2026-10-04 10:42 AEST ! fail GLM lab llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 lab_decode_c1_tps 10.4 tok/s · lab_prefill_tps 132 tok/s gates T21 2026-10-04 10:31 AEST ! fail GLM stability llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 decode_window_min_tps 9.86 tok/s · tool_calls 22 calls T21 2026-10-04 10:12 AEST ✓ pass GLM TTFT llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 ttft_prefill_tps 130 tok/s T21 2026-10-04 09:30 AEST ! fail GLM lab llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 lab_decode_c1_tps 10.5 tok/s · lab_prefill_tps 126 tok/s gates T21 2026-10-04 09:29 AEST ✓ pass GLM load llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 load_seconds 79.7 s T21 2026-10-04 09:17 AEST ✕ crash GLM lab llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 lab_decode_c1_tps unavailable · lab_prefill_tps unavailable harness T21 2026-10-04 09:07 AEST ! fail GLM stability llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 decode_window_min_tps 11.4 tok/s · tool_calls 26 calls T21 2026-10-04 08:51 AEST ✓ pass GLM TTFT llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 ttft_prefill_tps 166 tok/s T21 2026-10-04 08:27 AEST ! fail GLM lab llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 lab_decode_c1_tps 12.3 tok/s · lab_prefill_tps 158 tok/s gates T21 2026-10-04 08:25 AEST ✓ pass GLM load llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 load_seconds 101 s T21 2026-10-04 08:17 AEST ! fail Flash-Next stability llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012 decode_window_min_tps unavailable · tool_calls 0 calls T11 2026-10-04 08:06 AEST ✓ pass Flash-Next load llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012 load_seconds 39.9 s T11 2026-10-04 06:02 AEST ✕ crash GLM stability ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2 decode_window_min_tps unavailable · tool_calls 0 calls measure T22b 2026-10-04 05:48 AEST ✓ pass GLM TTFT ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2 ttft_prefill_tps 199 tok/s T22b 2026-10-04 05:33 AEST ! fail GLM lab ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2 lab_decode_c1_tps 8.60 tok/s · lab_prefill_tps 147 tok/s gates T22b 2026-10-04 05:32 AEST ✓ pass GLM load ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2 load_seconds 51.9 s T22b 2026-10-04 05:16 AEST ! fail GLM lab llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 lab_decode_c1_tps 9.50 tok/s · lab_prefill_tps 199 tok/s gates T20b 2026-10-04 05:06 AEST ! fail GLM stability llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 decode_window_min_tps 8.85 tok/s · tool_calls 30 calls T20b 2026-10-04 04:54 AEST ✓ pass GLM TTFT llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 ttft_prefill_tps 212 tok/s T20b 2026-10-04 04:26 AEST ! fail GLM lab llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 lab_decode_c1_tps 9.50 tok/s · lab_prefill_tps 200 tok/s gates T20b 2026-10-04 04:26 AEST ✓ pass GLM load llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 load_seconds 45.9 s T20b 2026-10-04 03:40 AEST ! fail GLM lab llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 lab_decode_c1_tps 8.40 tok/s · lab_prefill_tps 71 tok/s gates T20b 2026-10-04 03:30 AEST ! fail GLM stability llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 decode_window_min_tps 7.65 tok/s · tool_calls 25 calls T20b 2026-10-04 02:55 AEST ✓ pass GLM TTFT llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 ttft_prefill_tps 75.9 tok/s T20b 2026-10-04 02:26 AEST ! fail GLM lab llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 lab_decode_c1_tps 8.30 tok/s · lab_prefill_tps 65 tok/s gates T20b 2026-10-04 02:25 AEST ✓ pass GLM load llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 load_seconds 58.2 s T20b 2026-10-04 01:51 AEST ✕ crash GLM stability ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2 decode_window_min_tps unavailable · tool_calls 0 calls measure T22b 2026-10-04 01:37 AEST ✓ pass GLM TTFT ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2 ttft_prefill_tps 201 tok/s T22b 2026-10-04 01:19 AEST ! fail GLM lab ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2 lab_decode_c1_tps 8.20 tok/s · lab_prefill_tps 147 tok/s gates T22b 2026-10-04 01:18 AEST ✓ pass GLM load ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2 load_seconds 47.7 s T22b 2026-10-04 01:01 AEST ! fail GLM lab ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 lab_decode_c1_tps 9.50 tok/s · lab_prefill_tps 151 tok/s gates T22b 2026-10-04 00:51 AEST ! fail GLM stability ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 decode_window_min_tps 9.46 tok/s · tool_calls 0 calls T22b 2026-10-04 00:37 AEST ✓ pass GLM TTFT ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 ttft_prefill_tps 203 tok/s T22b 2026-10-04 00:22 AEST ! fail GLM lab ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 lab_decode_c1_tps 10.0 tok/s · lab_prefill_tps 151 tok/s gates T22b 2026-10-04 00:22 AEST ✓ pass GLM load ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 load_seconds 46.1 s T22b 2026-10-03 23:31 AEST ! fail GLM lab llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 lab_decode_c1_tps 9.30 tok/s · lab_prefill_tps 80 tok/s gates T20b 2026-10-03 23:21 AEST ! fail GLM stability llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 decode_window_min_tps 8.80 tok/s · tool_calls 19 calls T20b 2026-10-03 22:50 AEST ✓ pass GLM TTFT llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 ttft_prefill_tps 84.4 tok/s T20b 2026-10-03 21:59 AEST ! fail GLM lab llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 lab_decode_c1_tps 9.20 tok/s · lab_prefill_tps 78 tok/s gates T20b 2026-10-03 21:58 AEST ✓ pass GLM load llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 load_seconds 49.0 s T20b 2026-10-03 12:48 AEST ! fail GLM lab glm-llamacpp-iq2m-128k-1x-gpu2 lab_decode_c1_tps 9.60 tok/s · lab_prefill_tps 81 tok/s gates T20b 2026-10-03 12:38 AEST ! fail GLM stability glm-llamacpp-iq2m-128k-1x-gpu2 decode_window_min_tps 8.91 tok/s · tool_calls 24 calls T20b 2026-10-03 12:07 AEST ✓ pass GLM TTFT glm-llamacpp-iq2m-128k-1x-gpu2 ttft_prefill_tps 86.3 tok/s T20b 2026-10-03 11:41 AEST ! fail GLM lab glm-llamacpp-iq2m-128k-1x-gpu2 lab_decode_c1_tps 9.60 tok/s · lab_prefill_tps 81 tok/s gates T20b 2026-10-03 11:41 AEST ✓ pass GLM load glm-llamacpp-iq2m-128k-1x-gpu2 load_seconds 45.9 s T20b 2026-10-03 11:14 AEST ✓ pass Flash-Next lab fn-llamacpp-iq4xs-128k-1x-fit-gpu2 lab_decode_c1_tps 26.1 tok/s · lab_prefill_tps 157 tok/s T11 2026-10-03 11:03 AEST ! fail Flash-Next stability fn-llamacpp-iq4xs-128k-1x-fit-gpu2 decode_window_min_tps 23.8 tok/s · tool_calls 49 calls T11 2026-10-03 10:49 AEST ✓ pass Flash-Next TTFT fn-llamacpp-iq4xs-128k-1x-fit-gpu2 ttft_prefill_tps 269 tok/s T11 2026-10-03 10:29 AEST ✓ pass Flash-Next lab fn-llamacpp-iq4xs-128k-1x-fit-gpu2 lab_decode_c1_tps 26.5 tok/s · lab_prefill_tps 157 tok/s T11 2026-10-03 10:29 AEST ✓ pass Flash-Next load fn-llamacpp-iq4xs-128k-1x-fit-gpu2 load_seconds 27.8 s T11 2026-10-03 10:28 AEST ✕ crash Flash-Next load fn-llamacpp-iq4xs-128k-1x-gpu2 load_seconds unavailable load T11 2026-10-03 08:17 AEST ✓ pass Flash-Next lab fn-trellis-3.05-0xsero-main-x8-gpu2 lab_decode_c1_tps 55.3 tok/s · lab_prefill_tps 2,228 tok/s T10 2026-10-03 08:06 AEST ! fail Flash-Next stability fn-trellis-3.05-0xsero-main-x8-gpu2 decode_window_min_tps 38.1 tok/s · tool_calls 63 calls T10 2026-10-03 08:06 AEST ✓ pass Flash-Next quality panel fn-trellis-3.05-0xsero-main-x8-gpu2 top1_agreement 0.99234 fraction · mean_kl 0.0009667 nats T10 2026-10-03 08:04 AEST ✓ pass Flash-Next TTFT fn-trellis-3.05-0xsero-main-x8-gpu2 ttft_prefill_tps 2,263 tok/s T10 2026-10-03 07:48 AEST ✓ pass Flash-Next kit sweep fn-trellis-3.05-0xsero-main-x8-gpu2 decode_c1_tps 46.9 tok/s · prefill_32k_tps 2,433 tok/s T10 2026-10-03 07:43 AEST ✓ pass Flash-Next lab fn-trellis-3.05-0xsero-main-x8-gpu2 lab_decode_c1_tps 49.3 tok/s · lab_prefill_tps 2,176 tok/s T10 2026-10-03 07:40 AEST ✓ pass Flash-Next load fn-trellis-3.05-0xsero-main-x8-gpu2 load_seconds 179 s T10 2026-10-03 06:54 AEST ✓ pass none nvme nvme-rand-1tb-ngram rand_read_iops 168,600 IOPS · rand_read_gbps 3.35 GB/s T05a 2026-10-03 06:53 AEST ✓ pass none hw bw hw-bw-gpu0-current h2d_gbps 6.03 GB/s · d2h_gbps 6.60 GB/s T05a 2026-10-03 06:53 AEST ✓ pass none hw bw hw-bw-gpu2-current h2d_gbps 13.5 GB/s · d2h_gbps 13.2 GB/s T05a
Runs / none
20261002T205342Z-hw-bw-gpu2-current-hw_bw ✓ pass hw bw T05a eligibility: none
Config hw-bw-gpu2-current Started 2026-10-03 06:53 AEST Finished 2026-10-03 06:53 AEST (0.1 min) Exit status 0 Notes 1 GiB pinned transfers; full table in detail
Launch no container: static measurement
Metrics context.configured – context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run h2d_gbps 13.5 GB/s {"1MB_pin": 13.15, "1MB_page": 13.02, "4MB_pin": 13.39, "4MB_page": 13.14, "16MB_pin": 13.45, "16MB_page": 13.19, "64MB_pin": 13.46, "64MB_page": 13.18, "256MB_pin": 13.47, "256MB_page": 13.16, "1024MB_pin": 13.47, "1024MB_page": 13.11}d2h_gbps 13.2 GB/s {"1MB_pin": 12.98, "1MB_page": 7.19, "4MB_pin": 13.14, "4MB_page": 10.06, "16MB_pin": 13.2, "16MB_page": 11.35, "64MB_pin": 13.21, "64MB_page": 12.57, "256MB_pin": 13.21, "256MB_page": 12.61, "1024MB_pin": 13.22, "1024MB_page": 12.69}host_memcpy_gbps 16.6 GB/s
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 483c32c7f11d39f2f1e84c87602db3e166dc9fff cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/none/20261002T205342Z-hw-bw-gpu2-current-hw_bw.json
Runs / none
20261002T205350Z-hw-bw-gpu0-current-hw_bw ✓ pass hw bw T05a eligibility: none
Config hw-bw-gpu0-current Started 2026-10-03 06:53 AEST Finished 2026-10-03 06:54 AEST (0.2 min) Exit status 0 Notes 1 GiB pinned transfers; full table in detail
Launch no container: static measurement
Metrics context.configured – context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run h2d_gbps 6.03 GB/s {"1MB_pin": 6.48, "1MB_page": 6.08, "4MB_pin": 6.6, "4MB_page": 6.1, "16MB_pin": 6.57, "16MB_page": 6.11, "64MB_pin": 6.09, "64MB_page": 6.05, "256MB_pin": 6.03, "256MB_page": 6.01, "1024MB_pin": 6.03, "1024MB_page": 5.99}d2h_gbps 6.60 GB/s {"1MB_pin": 6.51, "1MB_page": 4.62, "4MB_pin": 6.58, "4MB_page": 5.88, "16MB_pin": 6.59, "16MB_page": 6.34, "64MB_pin": 6.6, "64MB_page": 6.48, "256MB_pin": 6.6, "256MB_page": 6.51, "1024MB_pin": 6.6, "1024MB_page": 6.52}host_memcpy_gbps 16.6 GB/s
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 483c32c7f11d39f2f1e84c87602db3e166dc9fff cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/none/20261002T205350Z-hw-bw-gpu0-current-hw_bw.json
Runs / none
20261002T205409Z-nvme-rand-1tb-ngram-nvme_rand ✓ pass nvme T05a eligibility: none
Config nvme-rand-1tb-ngram Started 2026-10-03 06:54 AEST Finished 2026-10-03 06:55 AEST (1.1 min) Exit status 0 Notes ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5/ngram_embedding.safetensors
Launch no container: static measurement
Metrics context.configured – context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run rand_read_iops 168,600 IOPS "16 KiB, 32 threads"rand_read_gbps 3.35 GB/s [{"bs": 4096, "threads": 1, "GBps": 0.075, "kIOPS": 18.2}, {"bs": 4096, "threads": 8, "GBps": 0.502, "kIOPS": 122.6}, {"bs": 4096, "threads": 32, "GBps": 0.714, "kIOPS": 174.4}, {"bs": 4096, "threads": 64, "GBps": 0.697, "kIOPS": 170.3}, {"bs": 16384, "threads": 1, "GBps": 0.2, "kIOPS": 12.2}, {"bs": 16384, "threads": 8, "GBps": 1.328, "kIOPS": 81.1}, {"bs": 16384, "threads": 32, "GBps": 2.763, "kIOPS": 168.6}, {"bs": 16384, "threads": 64, "GBps": 2.678, "kIOPS": 163.4}, {"bs": 65536, "threads": 1, "GBps": 0.47, "kIOPS": 7.2}, {"bs": 65536, "threads": 8, "GBps": 2.96, "kIOPS": 45.2}, {"bs": 65536, "threads": 32, "GBps": 3.35, "kIOPS": 51.1}, {"bs": 65536, "threads": 64, "GBps": 3.35, "kIOPS": 51.1}, {"bs": 1048576, "threads": 1, "GBps": 1.972, "kIOPS": 1.9}, {"bs": 1048576, "threads": 8, "GBps": 3.35, "kIOPS": 3.2}, {"bs": 1048576, "threads": 32, "GBps": 3.353, "kIOPS": 3.2}, {"bs": 1048576, "threads": 64, "GBps": 3.355, "kIOPS": 3.2}]
Cache, layout and host Cache warm Layout current · 0 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 483c32c7f11d39f2f1e84c87602db3e166dc9fff cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/none/20261002T205409Z-nvme-rand-1tb-ngram-nvme_rand.json
Runs / Qwen3.8-Flash-Next / fn-trellis-3.05-0xsero-main-x8-gpu2
20261002T214031Z-fn-trellis-3.05-0xsero-main-x8-gpu2-load ✓ pass load T10 eligibility: none
Config fn-trellis-3.05-0xsero-main-x8-gpu2 Started 2026-10-03 07:40 AEST Finished 2026-10-03 07:43 AEST (3.0 min) Exit status None Notes –
Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261002T214031Z-lab-fn-trellis-3.05-0xsero-main-x8-gpu2/docker-run.txt) Copy
docker run -d --name lab-fn-trellis-3.05-0xsero-main-x8-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3Digest sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/1e72b250cb529c6d81bf11b88740d9a0b22f62de Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb.json sha256 248d18e53c5004810964b493e2676bb465a9386a5c527a8b88e67cbfd0dffa68MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 179 s host_ram_drop_gb 63.7 GB vram_ready_mib 21,060 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 966c1355dde3774c7406f9d18e6b8468d4a10fdb cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261002T214031Z-fn-trellis-3.05-0xsero-main-x8-gpu2-load.json
Runs / Qwen3.8-Flash-Next / fn-trellis-3.05-0xsero-main-x8-gpu2
20261002T214330Z-fn-trellis-3.05-0xsero-main-x8-gpu2-lab ✓ pass lab T10 eligibility: none
Config fn-trellis-3.05-0xsero-main-x8-gpu2 Started 2026-10-03 07:43 AEST Finished 2026-10-03 07:48 AEST (5.3 min) Exit status 0 Notes served=flashnext
Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261002T214031Z-lab-fn-trellis-3.05-0xsero-main-x8-gpu2/docker-run.txt) Copy
docker run -d --name lab-fn-trellis-3.05-0xsero-main-x8-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3Digest sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/1e72b250cb529c6d81bf11b88740d9a0b22f62de Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb.json sha256 248d18e53c5004810964b493e2676bb465a9386a5c527a8b88e67cbfd0dffa68MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max 166,667 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}lab_decode_c1_tps 49.3 tok/s lab_prefill_tps 2,176 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 966c1355dde3774c7406f9d18e6b8468d4a10fdb cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261002T214330Z-fn-trellis-3.05-0xsero-main-x8-gpu2-lab.json
Runs / Qwen3.8-Flash-Next / fn-trellis-3.05-0xsero-main-x8-gpu2
20261002T214848Z-fn-trellis-3.05-0xsero-main-x8-gpu2-kit_sweep ✓ pass kit sweep T10 eligibility: none
Config fn-trellis-3.05-0xsero-main-x8-gpu2 Started 2026-10-03 07:48 AEST Finished 2026-10-03 08:04 AEST (15.9 min) Exit status 0 Notes sweep status: DONE
Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261002T214031Z-lab-fn-trellis-3.05-0xsero-main-x8-gpu2/docker-run.txt) Copy
docker run -d --name lab-fn-trellis-3.05-0xsero-main-x8-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3Digest sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/1e72b250cb529c6d81bf11b88740d9a0b22f62de Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb.json sha256 248d18e53c5004810964b493e2676bb465a9386a5c527a8b88e67cbfd0dffa68MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run prefill_8k_tps 2,292 tok/s {"min": 2277.9, "max": 2305.2, "n": 3}prefill_16k_tps 2,401 tok/s {"min": 2400.6, "max": 2403.1, "n": 3}prefill_32k_tps 2,433 tok/s {"min": 2431.6, "max": 2433.5, "n": 3}prefill_64k_tps 2,415 tok/s {"min": 2413.8, "max": 2415.0, "n": 3}decode_c1_tps 46.9 tok/s {"min": 45.1, "max": 48.7, "power_w": null}decode_c2_tps 51.6 tok/s {"min": 51.1, "max": 52.11, "power_w": null}decode_c3_tps 53.9 tok/s {"min": 51.82, "max": 55.98, "power_w": null}decode_c4_tps 51.8 tok/s {"min": 50.57, "max": 53.12, "power_w": null}decode_c1_32k_tps 48.5 tok/s {"min": 48.51, "max": 48.6, "power_w": null}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 966c1355dde3774c7406f9d18e6b8468d4a10fdb cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261002T214848Z-fn-trellis-3.05-0xsero-main-x8-gpu2-kit_sweep.json
Runs / Qwen3.8-Flash-Next / fn-trellis-3.05-0xsero-main-x8-gpu2
20261002T220441Z-fn-trellis-3.05-0xsero-main-x8-gpu2-ttft ✓ pass TTFT T10 eligibility: none
Config fn-trellis-3.05-0xsero-main-x8-gpu2 Started 2026-10-03 08:04 AEST Finished 2026-10-03 08:06 AEST (1.5 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261002T214031Z-lab-fn-trellis-3.05-0xsero-main-x8-gpu2/docker-run.txt) Copy
docker run -d --name lab-fn-trellis-3.05-0xsero-main-x8-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3Digest sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/1e72b250cb529c6d81bf11b88740d9a0b22f62de Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb.json sha256 248d18e53c5004810964b493e2676bb465a9386a5c527a8b88e67cbfd0dffa68MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 2,263 tok/s {"prompt_tokens": [13679, 13707, 13790], "ttft_s": [6.049, 6.056, 6.045]}ttft_prefill_32k_tps 2,368 tok/s {"prompt_tokens": [54780, 54797, 54620], "ttft_s": [23.132, 23.12, 23.139]}ttft_prefill_tps 2,263 tok/s {"prompt_tokens": [13679, 13707, 13790], "ttft_s": [6.049, 6.056, 6.045]}
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 966c1355dde3774c7406f9d18e6b8468d4a10fdb cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261002T220441Z-fn-trellis-3.05-0xsero-main-x8-gpu2-ttft.json
Runs / Qwen3.8-Flash-Next / fn-trellis-3.05-0xsero-main-x8-gpu2
20261002T220609Z-fn-trellis-3.05-0xsero-main-x8-gpu2-quality_panel ✓ pass quality panel T10 eligibility: none
Config fn-trellis-3.05-0xsero-main-x8-gpu2 Started 2026-10-03 08:06 AEST Finished 2026-10-03 08:06 AEST (0.4 min) Exit status 0 Notes panel=qwen3.8-flash-next-exl3-ref-panel.json rc=0
Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261002T214031Z-lab-fn-trellis-3.05-0xsero-main-x8-gpu2/docker-run.txt) Copy
docker run -d --name lab-fn-trellis-3.05-0xsero-main-x8-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3Digest sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/1e72b250cb529c6d81bf11b88740d9a0b22f62de Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb.json sha256 248d18e53c5004810964b493e2676bb465a9386a5c527a8b88e67cbfd0dffa68MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run top1_agreement 0.99234 fraction mean_kl 0.0009667 nats in_band yes "top1>=0.987, KL<=0.0013"
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 966c1355dde3774c7406f9d18e6b8468d4a10fdb cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261002T220609Z-fn-trellis-3.05-0xsero-main-x8-gpu2-quality_panel.json
Runs / Qwen3.8-Flash-Next / fn-trellis-3.05-0xsero-main-x8-gpu2
20261002T220633Z-fn-trellis-3.05-0xsero-main-x8-gpu2-stability ! fail stability T10 eligibility: none
Config fn-trellis-3.05-0xsero-main-x8-gpu2 Started 2026-10-03 08:06 AEST Finished 2026-10-03 08:17 AEST (10.5 min) Exit status 0 Notes stability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 63 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 38.077 < floor 50.0
Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261002T214031Z-lab-fn-trellis-3.05-0xsero-main-x8-gpu2/docker-run.txt) Copy
docker run -d --name lab-fn-trellis-3.05-0xsero-main-x8-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3Digest sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/1e72b250cb529c6d81bf11b88740d9a0b22f62de Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb.json sha256 248d18e53c5004810964b493e2676bb465a9386a5c527a8b88e67cbfd0dffa68MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max 30,034 tokens context.headroom_min 174,233 tokens duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 63 calls {"per_min": 6.3, "requests": 45}parse_failures 0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 0 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /v1/tokenize", "saved": "artifacts: proxy/repetition/"}decode_window_min_tps 38.1 tok/s window 1: 42.97 tok/s window 2: 43.5 tok/s window 3: 47.28 tok/s window 4: 41.43 tok/s window 5: 43.59 tok/s window 6: 42.67 tok/s min 41.43 · max 47.28 tok/s {"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 331, "min_at_generation_s": 63.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}decode_whole_run_tps 43.1 tok/s {"generation_seconds": 450.481, "generated_tokens": 19411.4}tasks_attempted 2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 1 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom no {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 208.0, "requests": 18, "tool_calls": 26, "completion_tokens": 6387, "occupied_max": 12502}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 392.0, "requests": 27, "tool_calls": 37, "completion_tokens": 13133, "occupied_max": 30034}] listtasks_over_64k [] instance idscompletion_tokens 19,520 tokens {"finish_reasons": {"tool_calls": 45}}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Agentic summary: 600 s, 63 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 38.1 tok/s, whole run 43.1 tok/s.
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 966c1355dde3774c7406f9d18e6b8468d4a10fdb cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261002T220633Z-fn-trellis-3.05-0xsero-main-x8-gpu2-stability.json
Runs / Qwen3.8-Flash-Next / fn-trellis-3.05-0xsero-main-x8-gpu2
20261002T221703Z-fn-trellis-3.05-0xsero-main-x8-gpu2-lab ✓ pass lab T10 eligibility: none
Config fn-trellis-3.05-0xsero-main-x8-gpu2 Started 2026-10-03 08:17 AEST Finished 2026-10-03 08:22 AEST (5.0 min) Exit status 0 Notes served=flashnext
Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261002T214031Z-lab-fn-trellis-3.05-0xsero-main-x8-gpu2/docker-run.txt) Copy
docker run -d --name lab-fn-trellis-3.05-0xsero-main-x8-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3Digest sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/1e72b250cb529c6d81bf11b88740d9a0b22f62de Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb.json sha256 248d18e53c5004810964b493e2676bb465a9386a5c527a8b88e67cbfd0dffa68MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max 166,667 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}lab_decode_c1_tps 55.3 tok/s lab_prefill_tps 2,228 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 966c1355dde3774c7406f9d18e6b8468d4a10fdb cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261002T221703Z-fn-trellis-3.05-0xsero-main-x8-gpu2-lab.json
Runs / Qwen3.8-Flash-Next / fn-llamacpp-iq4xs-128k-1x-gpu2
20261003T002815Z-fn-llamacpp-iq4xs-128k-1x-gpu2-load ✕ crash load T11 eligibility: none failed stage: load
Config fn-llamacpp-iq4xs-128k-1x-gpu2 Started 2026-10-03 10:28 AEST Finished 2026-10-03 10:28 AEST (0.4 min) Exit status 1 Notes user, abort
0.18.994.031 E ggml_backend_cuda_buffer_type_alloc_buffer: allocating 65358.17 MiB on device 0: cudaMalloc failed: out of memory
0.18.994.037 E alloc_tensor_range: failed to allocate CUDA0 buffer of size 68533006336
0.19.139.695 E llama_model_load: error loading model: unable to allocate CUDA0 buffer
0.19.139.702 E llama_model_load_from_file_impl: failed to load model
0.19.139.708 E cmn common_init_: failed to load model '/models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf'
0.19.139.713 E srv load_model: failed to load model, '/models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf'
0.19.139.715 I srv operator(): operator(): cleaning up before exit...
0.19.140.570 E srv llama_server: exiting due to model loading error
Launch recorded by tools/run.py (artifacts/runs/20261003T002815Z-lab-fn-llamacpp-iq4xs-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-fn-llamacpp-iq4xs-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 -ot per_layer_token_embd.weight=CPU Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-1x.json sha256 697275c75292e4ef45a02a99b2a98c23b7a55630ed0430af0972a1dcf84b47d8MTP / speculative off argv /app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 -ot per_layer_token_embd.weight=CPUenv CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/Qwen3.8-Flash-Next-GGUF @928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds unavailable not produced (crash/load) host_ram_drop_gb unavailable not produced (crash/load) vram_ready_mib unavailable not produced (crash/load)
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 1fa796f9d509705f24dc1870357869cf7c293bc2 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261003T002815Z-fn-llamacpp-iq4xs-128k-1x-gpu2-load.json
Runs / Qwen3.8-Flash-Next / fn-llamacpp-iq4xs-128k-1x-fit-gpu2
20261003T002901Z-fn-llamacpp-iq4xs-128k-1x-fit-gpu2-load ✓ pass load T11 eligibility: none
Config fn-llamacpp-iq4xs-128k-1x-fit-gpu2 Started 2026-10-03 10:29 AEST Finished 2026-10-03 10:29 AEST (0.5 min) Exit status None Notes –
Other runs of this config: load lab lab TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T002901Z-lab-fn-llamacpp-iq4xs-128k-1x-fit-gpu2/docker-run.txt) Copy
docker run -d --name lab-fn-llamacpp-iq4xs-128k-1x-fit-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-1x-fit.json sha256 67f1655622a5b08eae6fb8c89ca4a189aab6f7578a1a740baae4832ddc4e8160MTP / speculative off argv /app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/Qwen3.8-Flash-Next-GGUF @928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 27.8 s host_ram_drop_gb 2.30 GB vram_ready_mib 22,684 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 1fa796f9d509705f24dc1870357869cf7c293bc2 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261003T002901Z-fn-llamacpp-iq4xs-128k-1x-fit-gpu2-load.json
Runs / Qwen3.8-Flash-Next / fn-llamacpp-iq4xs-128k-1x-fit-gpu2
20261003T002929Z-fn-llamacpp-iq4xs-128k-1x-fit-gpu2-lab ✓ pass lab T11 eligibility: none
Config fn-llamacpp-iq4xs-128k-1x-fit-gpu2 Started 2026-10-03 10:29 AEST Finished 2026-10-03 10:49 AEST (19.9 min) Exit status 0 Notes served=/models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf
Other runs of this config: load lab lab TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T002901Z-lab-fn-llamacpp-iq4xs-128k-1x-fit-gpu2/docker-run.txt) Copy
docker run -d --name lab-fn-llamacpp-iq4xs-128k-1x-fit-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-1x-fit.json sha256 67f1655622a5b08eae6fb8c89ca4a189aab6f7578a1a740baae4832ddc4e8160MTP / speculative off argv /app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/Qwen3.8-Flash-Next-GGUF @928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max 106,294 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}lab_decode_c1_tps 26.5 tok/s lab_prefill_tps 157 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 1fa796f9d509705f24dc1870357869cf7c293bc2 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261003T002929Z-fn-llamacpp-iq4xs-128k-1x-fit-gpu2-lab.json
Runs / Qwen3.8-Flash-Next / fn-llamacpp-iq4xs-128k-1x-fit-gpu2
20261003T004923Z-fn-llamacpp-iq4xs-128k-1x-fit-gpu2-ttft ✓ pass TTFT T11 eligibility: none
Config fn-llamacpp-iq4xs-128k-1x-fit-gpu2 Started 2026-10-03 10:49 AEST Finished 2026-10-03 11:03 AEST (14.5 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab lab TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T002901Z-lab-fn-llamacpp-iq4xs-128k-1x-fit-gpu2/docker-run.txt) Copy
docker run -d --name lab-fn-llamacpp-iq4xs-128k-1x-fit-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-1x-fit.json sha256 67f1655622a5b08eae6fb8c89ca4a189aab6f7578a1a740baae4832ddc4e8160MTP / speculative off argv /app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/Qwen3.8-Flash-Next-GGUF @928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 269 tok/s {"prompt_tokens": [13827, 13626, 13781], "ttft_s": [50.985, 51.655, 51.198]}ttft_prefill_32k_tps 230 tok/s {"prompt_tokens": [54692, 54711, 54604], "ttft_s": [237.983, 237.379, 240.167]}ttft_prefill_tps 269 tok/s {"prompt_tokens": [13827, 13626, 13781], "ttft_s": [50.985, 51.655, 51.198]}
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 1fa796f9d509705f24dc1870357869cf7c293bc2 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261003T004923Z-fn-llamacpp-iq4xs-128k-1x-fit-gpu2-ttft.json
Runs / Qwen3.8-Flash-Next / fn-llamacpp-iq4xs-128k-1x-fit-gpu2
20261003T010352Z-fn-llamacpp-iq4xs-128k-1x-fit-gpu2-stability ! fail stability T11 eligibility: none
Config fn-llamacpp-iq4xs-128k-1x-fit-gpu2 Started 2026-10-03 11:03 AEST Finished 2026-10-03 11:14 AEST (10.5 min) Exit status 0 Notes stability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 49 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: 1 repetition hit(s); decode window min 23.83 < floor 50.0
Other runs of this config: load lab lab TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T002901Z-lab-fn-llamacpp-iq4xs-128k-1x-fit-gpu2/docker-run.txt) Copy
docker run -d --name lab-fn-llamacpp-iq4xs-128k-1x-fit-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-1x-fit.json sha256 67f1655622a5b08eae6fb8c89ca4a189aab6f7578a1a740baae4832ddc4e8160MTP / speculative off argv /app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/Qwen3.8-Flash-Next-GGUF @928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max 16,819 tokens context.headroom_min 114,087 tokens duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 49 calls {"per_min": 4.9, "requests": 32}parse_failures 0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 0 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 1 hits {"ngram64x3": 1, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}decode_window_min_tps 23.8 tok/s window 1: 24.81 tok/s window 2: 24.06 tok/s window 3: 25.42 tok/s window 4: 24.9 tok/s window 5: 24.6 tok/s min 24.06 · max 25.42 tok/s {"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 271, "min_at_generation_s": 105.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}decode_whole_run_tps 24.8 tok/s {"generation_seconds": 390.845, "generated_tokens": 9687.8}tasks_attempted 2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 1 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom no {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 306.9, "requests": 17, "tool_calls": 26, "completion_tokens": 4940, "occupied_max": 16819}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 293.1, "requests": 16, "tool_calls": 24, "completion_tokens": 4788, "occupied_max": 12213}] listtasks_over_64k [] instance idscompletion_tokens 9,728 tokens {"finish_reasons": {"tool_calls": 32}}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Agentic summary: 600 s, 49 tool calls, 0 parse failures, 1 repetition hits, tasks solved 1/2, decode min window 23.8 tok/s, whole run 24.8 tok/s.
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 1fa796f9d509705f24dc1870357869cf7c293bc2 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261003T010352Z-fn-llamacpp-iq4xs-128k-1x-fit-gpu2-stability.json
Runs / Qwen3.8-Flash-Next / fn-llamacpp-iq4xs-128k-1x-fit-gpu2
20261003T011420Z-fn-llamacpp-iq4xs-128k-1x-fit-gpu2-lab ✓ pass lab T11 eligibility: none
Config fn-llamacpp-iq4xs-128k-1x-fit-gpu2 Started 2026-10-03 11:14 AEST Finished 2026-10-03 11:40 AEST (26.5 min) Exit status 0 Notes served=/models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf
Other runs of this config: load lab lab TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T002901Z-lab-fn-llamacpp-iq4xs-128k-1x-fit-gpu2/docker-run.txt) Copy
docker run -d --name lab-fn-llamacpp-iq4xs-128k-1x-fit-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-1x-fit.json sha256 67f1655622a5b08eae6fb8c89ca4a189aab6f7578a1a740baae4832ddc4e8160MTP / speculative off argv /app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/Qwen3.8-Flash-Next-GGUF @928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max 106,294 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}lab_decode_c1_tps 26.1 tok/s lab_prefill_tps 157 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 1fa796f9d509705f24dc1870357869cf7c293bc2 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261003T011420Z-fn-llamacpp-iq4xs-128k-1x-fit-gpu2-lab.json
Runs / GLM-5.3-Flash / glm-llamacpp-iq2m-128k-1x-gpu2
20261003T014103Z-glm-llamacpp-iq2m-128k-1x-gpu2-load ✓ pass load T20b eligibility: none
Config glm-llamacpp-iq2m-128k-1x-gpu2 Started 2026-10-03 11:41 AEST Finished 2026-10-03 11:41 AEST (0.8 min) Exit status None Notes –
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T014103Z-lab-glm-llamacpp-iq2m-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-glm-llamacpp-iq2m-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x.json sha256 01715887da6d653bd4cc63d959d62744ec8f44f30ca1a60627052180981e3fc1MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 45.9 s host_ram_drop_gb 2.23 GB vram_ready_mib 22,784 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 4b12eecb28cf351fc09677fc509debe454394308 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T014103Z-glm-llamacpp-iq2m-128k-1x-gpu2-load.json
Runs / GLM-5.3-Flash / glm-llamacpp-iq2m-128k-1x-gpu2
20261003T014149Z-glm-llamacpp-iq2m-128k-1x-gpu2-lab ! fail lab T20b eligibility: none failed stage: gates
Config glm-llamacpp-iq2m-128k-1x-gpu2 Started 2026-10-03 11:41 AEST Finished 2026-10-03 12:07 AEST (26.0 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T014103Z-lab-glm-llamacpp-iq2m-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-glm-llamacpp-iq2m-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x.json sha256 01715887da6d653bd4cc63d959d62744ec8f44f30ca1a60627052180981e3fc1MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 9.60 tok/s lab_prefill_tps 81 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 4b12eecb28cf351fc09677fc509debe454394308 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T014149Z-glm-llamacpp-iq2m-128k-1x-gpu2-lab.json
Runs / GLM-5.3-Flash / glm-llamacpp-iq2m-128k-1x-gpu2
20261003T020751Z-glm-llamacpp-iq2m-128k-1x-gpu2-ttft ✓ pass TTFT T20b eligibility: none
Config glm-llamacpp-iq2m-128k-1x-gpu2 Started 2026-10-03 12:07 AEST Finished 2026-10-03 12:38 AEST (30.5 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T014103Z-lab-glm-llamacpp-iq2m-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-glm-llamacpp-iq2m-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x.json sha256 01715887da6d653bd4cc63d959d62744ec8f44f30ca1a60627052180981e3fc1MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 86.3 tok/s {"prompt_tokens": [10742, 10816, 10721], "ttft_s": [124.451, 126.526, 123.598]}ttft_prefill_32k_tps 88.5 tok/s {"prompt_tokens": [42736, 42800, 42830], "ttft_s": [485.781, 483.408, 483.895]}ttft_prefill_tps 86.3 tok/s {"prompt_tokens": [10742, 10816, 10721], "ttft_s": [124.451, 126.526, 123.598]}
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit bff973e82a9b895daba523707af9c56454e76a39 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T020751Z-glm-llamacpp-iq2m-128k-1x-gpu2-ttft.json
Runs / GLM-5.3-Flash / glm-llamacpp-iq2m-128k-1x-gpu2
20261003T023819Z-glm-llamacpp-iq2m-128k-1x-gpu2-stability ! fail stability T20b eligibility: none
Config glm-llamacpp-iq2m-128k-1x-gpu2 Started 2026-10-03 12:38 AEST Finished 2026-10-03 12:48 AEST (10.1 min) Exit status 0 Notes stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 24 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 8.91 < floor 15.0
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T014103Z-lab-glm-llamacpp-iq2m-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-glm-llamacpp-iq2m-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x.json sha256 01715887da6d653bd4cc63d959d62744ec8f44f30ca1a60627052180981e3fc1MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 9,047 tokens context.headroom_min 121,785 tokens duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 24 calls {"per_min": 2.4, "requests": 13}parse_failures 0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 0 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}decode_window_min_tps 8.91 tok/s window 1: 8.98 tok/s window 2: 9.23 tok/s window 3: 9.09 tok/s window 4: 9.18 tok/s window 5: 9.04 tok/s window 6: 9.14 tok/s min 8.98 · max 9.23 tok/s {"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 328, "min_at_generation_s": 90.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}decode_whole_run_tps 9.14 tok/s {"generation_seconds": 447.593, "generated_tokens": 4091.5}tasks_attempted 1 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 0 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom no {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 600.0, "requests": 14, "tool_calls": 24, "completion_tokens": 4105, "occupied_max": 9047}] listtasks_over_64k [] instance idscompletion_tokens 4,105 tokens {"finish_reasons": {"tool_calls": 13}}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Agentic summary: 600 s, 24 tool calls, 0 parse failures, 0 repetition hits, tasks solved 0/1, decode min window 8.91 tok/s, whole run 9.14 tok/s.
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit bff973e82a9b895daba523707af9c56454e76a39 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T023819Z-glm-llamacpp-iq2m-128k-1x-gpu2-stability.json
Runs / GLM-5.3-Flash / glm-llamacpp-iq2m-128k-1x-gpu2
20261003T024824Z-glm-llamacpp-iq2m-128k-1x-gpu2-lab ! fail lab T20b eligibility: none failed stage: gates
Config glm-llamacpp-iq2m-128k-1x-gpu2 Started 2026-10-03 12:48 AEST Finished 2026-10-03 13:14 AEST (25.9 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T014103Z-lab-glm-llamacpp-iq2m-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-glm-llamacpp-iq2m-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x.json sha256 01715887da6d653bd4cc63d959d62744ec8f44f30ca1a60627052180981e3fc1MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 9.60 tok/s lab_prefill_tps 81 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit bff973e82a9b895daba523707af9c56454e76a39 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T024824Z-glm-llamacpp-iq2m-128k-1x-gpu2-lab.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2
20261003T115813Z-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2-load ✓ pass load T20b eligibility: none
Config llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 Started 2026-10-03 21:58 AEST Finished 2026-10-03 21:59 AEST (0.8 min) Exit status None Notes –
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T115813Z-lab-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-Q2_K:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-q2-k-128k-1x.json sha256 f7ab285db47877ab680a678ead1d3de7960b1a54121d89c804a51dc6c2477de0MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-Q2_K.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 49.0 s host_ram_drop_gb 2.22 GB vram_ready_mib 22,626 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 73bf82fd3254285cf46014dbef7e02e6ef6db9b3 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T115813Z-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2-load.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2
20261003T115902Z-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2-lab ! fail lab T20b eligibility: none failed stage: gates
Config llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 Started 2026-10-03 21:59 AEST Finished 2026-10-03 22:50 AEST (51.0 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T115813Z-lab-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-Q2_K:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-q2-k-128k-1x.json sha256 f7ab285db47877ab680a678ead1d3de7960b1a54121d89c804a51dc6c2477de0MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-Q2_K.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 9.20 tok/s lab_prefill_tps 78 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit bc100ca481d268577db82a7adcd92133db2241e4 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T115902Z-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2-lab.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2
20261003T125002Z-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2-ttft ✓ pass TTFT T20b eligibility: none
Config llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 Started 2026-10-03 22:50 AEST Finished 2026-10-03 23:21 AEST (31.5 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T115813Z-lab-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-Q2_K:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-q2-k-128k-1x.json sha256 f7ab285db47877ab680a678ead1d3de7960b1a54121d89c804a51dc6c2477de0MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-Q2_K.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 84.4 tok/s {"prompt_tokens": [10644, 10721, 10688], "ttft_s": [126.044, 127.442, 125.735]}ttft_prefill_32k_tps 84.9 tok/s {"prompt_tokens": [42901, 42759, 42872], "ttft_s": [504.114, 504.375, 505.138]}ttft_prefill_tps 84.4 tok/s {"prompt_tokens": [10644, 10721, 10688], "ttft_s": [126.044, 127.442, 125.735]}
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit bc100ca481d268577db82a7adcd92133db2241e4 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T125002Z-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2-ttft.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2
20261003T132135Z-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2-stability ! fail stability T20b eligibility: none
Config llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 Started 2026-10-03 23:21 AEST Finished 2026-10-03 23:31 AEST (10.1 min) Exit status 0 Notes stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 19 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 8.8 < floor 15.0
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T115813Z-lab-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-Q2_K:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-q2-k-128k-1x.json sha256 f7ab285db47877ab680a678ead1d3de7960b1a54121d89c804a51dc6c2477de0MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-Q2_K.json
Metrics context.configured 131,072 tokens context.occupied_max 6,658 tokens context.headroom_min 123,658 tokens duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 19 calls {"per_min": 1.9, "requests": 11}parse_failures 0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 0 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}decode_window_min_tps 8.80 tok/s window 1: 8.89 tok/s window 2: 8.87 tok/s window 3: 8.93 tok/s window 4: 8.88 tok/s window 5: 9.0 tok/s min 8.87 · max 9 tok/s {"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 267, "min_at_generation_s": 131.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}decode_whole_run_tps 8.93 tok/s {"generation_seconds": 386.503, "generated_tokens": 3451.5}tasks_attempted 1 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 0 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom no {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 600.0, "requests": 12, "tool_calls": 20, "completion_tokens": 3463, "occupied_max": 6658}] listtasks_over_64k [] instance idscompletion_tokens 3,463 tokens {"finish_reasons": {"tool_calls": 11}}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Agentic summary: 600 s, 19 tool calls, 0 parse failures, 0 repetition hits, tasks solved 0/1, decode min window 8.80 tok/s, whole run 8.93 tok/s.
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit bc100ca481d268577db82a7adcd92133db2241e4 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T132135Z-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2-stability.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2
20261003T133139Z-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2-lab ! fail lab T20b eligibility: none failed stage: gates
Config llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 Started 2026-10-03 23:31 AEST Finished 2026-10-04 00:21 AEST (50.3 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T115813Z-lab-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-Q2_K:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-q2-k-128k-1x.json sha256 f7ab285db47877ab680a678ead1d3de7960b1a54121d89c804a51dc6c2477de0MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-Q2_K.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 9.30 tok/s lab_prefill_tps 80 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit bc100ca481d268577db82a7adcd92133db2241e4 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T133139Z-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2-lab.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2
20261003T142203Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2-load ✓ pass load T22b eligibility: none
Config ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 Started 2026-10-04 00:22 AEST Finished 2026-10-04 00:22 AEST (0.8 min) Exit status None Notes –
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T142203Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-1x.json sha256 168e8b490c871dcebe6afed1dfd973be9985379eb748e37f17be98b86a86e44eMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 46.1 s host_ram_drop_gb 111 GB vram_ready_mib 14,936 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 57320e8fbf776bca99fc63e2de1647d22ef83f06 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T142203Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2-load.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2
20261003T142250Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2-lab ! fail lab T22b eligibility: none failed stage: gates
Config ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 Started 2026-10-04 00:22 AEST Finished 2026-10-04 00:37 AEST (14.7 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T142203Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-1x.json sha256 168e8b490c871dcebe6afed1dfd973be9985379eb748e37f17be98b86a86e44eMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": false, "context": true, "speed": false}lab_decode_c1_tps 10.0 tok/s lab_prefill_tps 151 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 57320e8fbf776bca99fc63e2de1647d22ef83f06 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T142250Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2-lab.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2
20261003T143733Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2-ttft ✓ pass TTFT T22b eligibility: none
Config ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 Started 2026-10-04 00:37 AEST Finished 2026-10-04 00:51 AEST (13.5 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T142203Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-1x.json sha256 168e8b490c871dcebe6afed1dfd973be9985379eb748e37f17be98b86a86e44eMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 203 tok/s {"prompt_tokens": [10751, 10785, 10716], "ttft_s": [52.615, 53.39, 52.749]}ttft_prefill_32k_tps 197 tok/s {"prompt_tokens": [42739, 42927, 42933], "ttft_s": [217.584, 218.259, 218.191]}ttft_prefill_tps 203 tok/s {"prompt_tokens": [10751, 10785, 10716], "ttft_s": [52.615, 53.39, 52.749]}
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 57320e8fbf776bca99fc63e2de1647d22ef83f06 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T143733Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2-ttft.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2
20261003T145106Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2-stability ! fail stability T22b eligibility: none
Config ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 Started 2026-10-04 00:51 AEST Finished 2026-10-04 01:01 AEST (10.1 min) Exit status 0 Notes stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: 47 tool-call parse failure(s); decode window min 9.458 < floor 15.0
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T142203Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-1x.json sha256 168e8b490c871dcebe6afed1dfd973be9985379eb748e37f17be98b86a86e44eMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 1,874 tokens context.headroom_min 129,125 tokens duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 0 calls {"per_min": 0.0, "requests": 47}parse_failures 47 responses {"wire": {"unparsed_markup": 47}, "harness_format_errors": {"no_tool_call": 47}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 47 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}decode_window_min_tps 9.46 tok/s window 1: 9.49 tok/s window 2: 9.55 tok/s window 3: 9.5 tok/s window 4: 9.46 tok/s window 5: 9.55 tok/s min 9.46 · max 9.55 tok/s {"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 277, "min_at_generation_s": 144.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}decode_whole_run_tps 9.50 tok/s {"generation_seconds": 396.975, "generated_tokens": 3772.8}tasks_attempted 16 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 0 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom no {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 51.1, "requests": 3, "tool_calls": 0, "completion_tokens": 310, "occupied_max": 1559}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 40.8, "requests": 3, "tool_calls": 0, "completion_tokens": 227, "occupied_max": 1641}, {"attempt": 3, "round": 1, "instance_id": "psf__requests-2317", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 39.7, "requests": 3, "tool_calls": 0, "completion_tokens": 220, "occupied_max": 1609}, {"attempt": 4, "round": 1, "instance_id": "matplotlib__matplotlib-23299", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 43.5, "requests": 3, "tool_calls": 0, "completion_tokens": 251, "occupied_max": 1874}, {"attempt": 5, "round": 1, "instance_id": "scikit-learn__scikit-learn-13439", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 41.3, "requests": 3, "tool_calls": 0, "completion_tokens": 233, "occupied_max": 1770}, {"attempt": 6, "round": 1, "instance_id": "astropy__astropy-12907", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 34.9, "requests": 3, "tool_calls": 0, "completion_tokens": 174, "occupied_max": 1716}, {"attempt": 7, "round": 1, "instance_id": "pytest-dev__pytest-7373", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 41.1, "requests": 3, "tool_calls": 0, "completion_tokens": 231, "occupied_max": 1641}, {"attempt": 8, "round": 1, "instance_id": "sympy__sympy-20590", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 52.5, "requests": 3, "tool_calls": 0, "completion_tokens": 339, "occupied_max": 1578}, {"attempt": 9, "round": 2, "instance_id": "django__django-11099", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 31.5, "requests": 3, "tool_calls": 0, "completion_tokens": 230, "occupied_max": 1559}, {"attempt": 10, "round": 2, "instance_id": "sympy__sympy-21612", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 35.7, "requests": 3, "tool_calls": 0, "completion_tokens": 270, "occupied_max": 1641}, {"attempt": 11, "round": 2, "instance_id": "psf__requests-2317", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 32.8, "requests": 3, "tool_calls": 0, "completion_tokens": 243, "occupied_max": 1609}, {"attempt": 12, "round": 2, "instance_id": "matplotlib__matplotlib-23299", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 33.5, "requests": 3, "tool_calls": 0, "completion_tokens": 249, "occupied_max": 1874}, {"attempt": 13, "round": 2, "instance_id": "scikit-learn__scikit-learn-13439", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 33.7, "requests": 3, "tool_calls": 0, "completion_tokens": 251, "occupied_max": 1770}, {"attempt": 14, "round": 2, "instance_id": "astropy__astropy-12907", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 25.9, "requests": 3, "tool_calls": 0, "completion_tokens": 178, "occupied_max": 1716}, {"attempt": 15, "round": 2, "instance_id": "pytest-dev__pytest-7373", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 30.0, "requests": 3, "tool_calls": 0, "completion_tokens": 217, "occupied_max": 1641}, {"attempt": 16, "round": 2, "instance_id": "sympy__sympy-20590", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 32.0, "requests": 3, "tool_calls": 0, "completion_tokens": 198, "occupied_max": 1459}] listtasks_over_64k [] instance idscompletion_tokens 3,821 tokens {"finish_reasons": {"stop": 47}}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Agentic summary: 600 s, 0 tool calls, 47 parse failures, 0 repetition hits, tasks solved 0/16, decode min window 9.46 tok/s, whole run 9.50 tok/s.
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 57320e8fbf776bca99fc63e2de1647d22ef83f06 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T145106Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2-stability.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2
20261003T150111Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2-lab ! fail lab T22b eligibility: none failed stage: gates
Config ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 Started 2026-10-04 01:01 AEST Finished 2026-10-04 01:18 AEST (17.7 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T142203Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-1x.json sha256 168e8b490c871dcebe6afed1dfd973be9985379eb748e37f17be98b86a86e44eMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": false, "context": true, "speed": false}lab_decode_c1_tps 9.50 tok/s lab_prefill_tps 151 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 57320e8fbf776bca99fc63e2de1647d22ef83f06 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T150111Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2-lab.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2
20261003T151858Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2-load ✓ pass load T22b eligibility: none
Config ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2 Started 2026-10-04 01:18 AEST Finished 2026-10-04 01:19 AEST (0.8 min) Exit status None Notes –
Other runs of this config: load lab ! TTFT stability ✕
Launch recorded by tools/run.py (artifacts/runs/20261003T151858Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp.json sha256 c2df636d0f6ab990abe7dd3c658b064f3c8df59d819b7fe907ce82a12b84a9c4MTP / speculative {"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 47.7 s host_ram_drop_gb 115 GB vram_ready_mib 22,318 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 664783f004ff58b6fcabb0691d6b1c8298e2e818 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T151858Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2-load.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2
20261003T151947Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2-lab ! fail lab T22b eligibility: none failed stage: gates
Config ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2 Started 2026-10-04 01:19 AEST Finished 2026-10-04 01:37 AEST (17.7 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
Other runs of this config: load lab ! TTFT stability ✕
Launch recorded by tools/run.py (artifacts/runs/20261003T151858Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp.json sha256 c2df636d0f6ab990abe7dd3c658b064f3c8df59d819b7fe907ce82a12b84a9c4MTP / speculative {"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": false, "context": true, "speed": false}lab_decode_c1_tps 8.20 tok/s lab_prefill_tps 147 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps unavailable not produced (fail/gates) mtp_acceptance_rate unavailable not produced (fail/gates)
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 937eb11054a0e7abd1b0e970f25ea2fda3f4cff9 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T151947Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2-lab.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2
20261003T153729Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2-ttft ✓ pass TTFT T22b eligibility: none
Config ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2 Started 2026-10-04 01:37 AEST Finished 2026-10-04 01:51 AEST (13.9 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab ! TTFT stability ✕
Launch recorded by tools/run.py (artifacts/runs/20261003T151858Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp.json sha256 c2df636d0f6ab990abe7dd3c658b064f3c8df59d819b7fe907ce82a12b84a9c4MTP / speculative {"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 201 tok/s {"prompt_tokens": [10740, 10677, 10728], "ttft_s": [53.545, 52.784, 53.644]}ttft_prefill_32k_tps 190 tok/s {"prompt_tokens": [42815, 42706, 42534], "ttft_s": [223.459, 224.655, 228.199]}ttft_prefill_tps 201 tok/s {"prompt_tokens": [10740, 10677, 10728], "ttft_s": [53.545, 52.784, 53.644]}
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 937eb11054a0e7abd1b0e970f25ea2fda3f4cff9 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T153729Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2-ttft.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2
20261003T155126Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2-stability ✕ crash stability T22b eligibility: none failed stage: measure
Config ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2 Started 2026-10-04 01:51 AEST Finished 2026-10-04 02:01 AEST (10.1 min) Exit status 0 Notes stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: engine connection failures (crash/OOM); 9 tool-call parse failure(s); decode floor not verifiable (too little generation)
Other runs of this config: load lab ! TTFT stability ✕
Launch recorded by tools/run.py (artifacts/runs/20261003T151858Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp.json sha256 c2df636d0f6ab990abe7dd3c658b064f3c8df59d819b7fe907ce82a12b84a9c4MTP / speculative {"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 1,641 tokens context.headroom_min 129,356 tokens duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 0 calls {"per_min": 0.0, "requests": 9}parse_failures 9 responses {"wire": {"unparsed_markup": 9}, "harness_format_errors": {"no_tool_call": 9}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 9 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}decode_window_min_tps unavailable only 57.72 s of generation; need > 120 s for one window decode_whole_run_tps 12.6 tok/s {"generation_seconds": 57.72, "generated_tokens": 729.8}tasks_attempted 5 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 0 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom yes {"upstream_connection_failures": 54, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 47.1, "requests": 3, "tool_calls": 0, "completion_tokens": 274, "occupied_max": 1559}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 36.4, "requests": 3, "tool_calls": 0, "completion_tokens": 248, "occupied_max": 1641}, {"attempt": 3, "round": 1, "instance_id": "psf__requests-2317", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 33.8, "requests": 3, "tool_calls": 0, "completion_tokens": 217, "occupied_max": 1609}, {"attempt": 4, "round": 1, "instance_id": "matplotlib__matplotlib-23299", "exit_status": "Error:BadGatewayError: litellm.BadGatewayError: BadGatewayError: OpenAIException - Cannot connect to host 127.0.0.1:18080 ssl:default [Connect call failed ('127.0.0.1', 18080)]", "submitted": false, "resolved": false, "seconds": 343.5}, {"attempt": 5, "round": 1, "instance_id": "scikit-learn__scikit-learn-13439", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 139.3}] listtasks_over_64k [] instance idscompletion_tokens 739 tokens {"finish_reasons": {"stop": 9}}mtp_accepted_tps unavailable not produced (crash/measure) mtp_acceptance_rate unavailable not produced (crash/measure)
Agentic summary: 600 s, 0 tool calls, 9 parse failures, 0 repetition hits, tasks solved 0/5, decode min window – tok/s, whole run 12.6 tok/s.
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 937eb11054a0e7abd1b0e970f25ea2fda3f4cff9 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T155126Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2-stability.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2
20261003T162532Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2-load ✓ pass load T20b eligibility: none
Config llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 Started 2026-10-04 02:25 AEST Finished 2026-10-04 02:26 AEST (1.0 min) Exit status None Notes –
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T162532Z-lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ3_XXS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x.json sha256 76879bccb4223161305f6a73f10b3a4f79db3609c18c993495ef170aa8ae9db8MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ3_XXS.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 58.2 s host_ram_drop_gb 1.89 GB vram_ready_mib 22,596 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 9fc82ec3597413063d86252387463dee533769ed cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T162532Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2-load.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2
20261003T162631Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2-lab ! fail lab T20b eligibility: none failed stage: gates
Config llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 Started 2026-10-04 02:26 AEST Finished 2026-10-04 02:55 AEST (28.7 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T162532Z-lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ3_XXS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x.json sha256 76879bccb4223161305f6a73f10b3a4f79db3609c18c993495ef170aa8ae9db8MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ3_XXS.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 8.30 tok/s lab_prefill_tps 65 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 9fc82ec3597413063d86252387463dee533769ed cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T162631Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2-lab.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2
20261003T165514Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2-ttft ✓ pass TTFT T20b eligibility: none
Config llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 Started 2026-10-04 02:55 AEST Finished 2026-10-04 03:30 AEST (34.8 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T162532Z-lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ3_XXS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x.json sha256 76879bccb4223161305f6a73f10b3a4f79db3609c18c993495ef170aa8ae9db8MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ3_XXS.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 75.9 tok/s {"prompt_tokens": [10736, 10754, 10667], "ttft_s": [140.833, 141.619, 141.761]}ttft_prefill_32k_tps 77.3 tok/s {"prompt_tokens": [42840, 42803, 42831], "ttft_s": [554.48, 553.698, 555.603]}ttft_prefill_tps 75.9 tok/s {"prompt_tokens": [10736, 10754, 10667], "ttft_s": [140.833, 141.619, 141.761]}
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 9fc82ec3597413063d86252387463dee533769ed cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T165514Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2-ttft.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2
20261003T173003Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2-stability ! fail stability T20b eligibility: none
Config llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 Started 2026-10-04 03:30 AEST Finished 2026-10-04 03:40 AEST (10.5 min) Exit status 0 Notes stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 25 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 7.646 < floor 15.0
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T162532Z-lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ3_XXS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x.json sha256 76879bccb4223161305f6a73f10b3a4f79db3609c18c993495ef170aa8ae9db8MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ3_XXS.json
Metrics context.configured 131,072 tokens context.occupied_max 4,875 tokens context.headroom_min 126,123 tokens duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 25 calls {"per_min": 2.5, "requests": 15}parse_failures 0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 0 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}decode_window_min_tps 7.65 tok/s window 1: 8.08 tok/s window 2: 7.78 tok/s window 3: 7.99 tok/s window 4: 8.16 tok/s window 5: 8.19 tok/s min 7.78 · max 8.19 tok/s {"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 269, "min_at_generation_s": 156.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}decode_whole_run_tps 8.08 tok/s {"generation_seconds": 388.925, "generated_tokens": 3143.9}tasks_attempted 2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 1 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom no {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 377.0, "requests": 10, "tool_calls": 17, "completion_tokens": 2133, "occupied_max": 4875}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 223.0, "requests": 6, "tool_calls": 9, "completion_tokens": 1027, "occupied_max": 4558}] listtasks_over_64k [] instance idscompletion_tokens 3,160 tokens {"finish_reasons": {"tool_calls": 15}}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Agentic summary: 600 s, 25 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 7.65 tok/s, whole run 8.08 tok/s.
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 9fc82ec3597413063d86252387463dee533769ed cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T173003Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2-stability.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2
20261003T174033Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2-lab ! fail lab T20b eligibility: none failed stage: gates
Config llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 Started 2026-10-04 03:40 AEST Finished 2026-10-04 04:25 AEST (45.4 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T162532Z-lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ3_XXS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x.json sha256 76879bccb4223161305f6a73f10b3a4f79db3609c18c993495ef170aa8ae9db8MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ3_XXS.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 8.40 tok/s lab_prefill_tps 71 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 9fc82ec3597413063d86252387463dee533769ed cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T174033Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2-lab.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2
20261003T182602Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2-load ✓ pass load T20b eligibility: none
Config llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 Started 2026-10-04 04:26 AEST Finished 2026-10-04 04:26 AEST (0.8 min) Exit status None Notes –
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T182602Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048.json sha256 89ee35a7fb8e6ec8fd63aa7e9e07a5ed1c3448eed3d10cebf357887e810a94feMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 45.9 s host_ram_drop_gb 3.54 GB vram_ready_mib 21,944 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit d4f093c724ebfd683feed823be35a11d4c009375 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T182602Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2-load.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2
20261003T182648Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2-lab ! fail lab T20b eligibility: none failed stage: gates
Config llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 Started 2026-10-04 04:26 AEST Finished 2026-10-04 04:54 AEST (27.5 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T182602Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048.json sha256 89ee35a7fb8e6ec8fd63aa7e9e07a5ed1c3448eed3d10cebf357887e810a94feMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 9.50 tok/s lab_prefill_tps 200 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit d4f093c724ebfd683feed823be35a11d4c009375 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T182648Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2-lab.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2
20261003T185415Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2-ttft ✓ pass TTFT T20b eligibility: none
Config llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 Started 2026-10-04 04:54 AEST Finished 2026-10-04 05:06 AEST (12.2 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T182602Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048.json sha256 89ee35a7fb8e6ec8fd63aa7e9e07a5ed1c3448eed3d10cebf357887e810a94feMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 212 tok/s {"prompt_tokens": [10787, 10703, 10755], "ttft_s": [50.583, 51.91, 50.672]}ttft_prefill_32k_tps 221 tok/s {"prompt_tokens": [42782, 42887, 42759], "ttft_s": [192.559, 193.833, 193.907]}ttft_prefill_tps 212 tok/s {"prompt_tokens": [10787, 10703, 10755], "ttft_s": [50.583, 51.91, 50.672]}
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit d4f093c724ebfd683feed823be35a11d4c009375 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T185415Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2-ttft.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2
20261003T190629Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2-stability ! fail stability T20b eligibility: none
Config llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 Started 2026-10-04 05:06 AEST Finished 2026-10-04 05:16 AEST (10.4 min) Exit status 0 Notes stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 30 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 8.854 < floor 15.0
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T182602Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048.json sha256 89ee35a7fb8e6ec8fd63aa7e9e07a5ed1c3448eed3d10cebf357887e810a94feMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 9,732 tokens context.headroom_min 121,298 tokens duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 30 calls {"per_min": 3.0, "requests": 21}parse_failures 0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 0 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}decode_window_min_tps 8.85 tok/s window 1: 9.19 tok/s window 2: 9.0 tok/s window 3: 8.95 tok/s window 4: 8.92 tok/s window 5: 9.14 tok/s min 8.92 · max 9.19 tok/s {"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 274, "min_at_generation_s": 232.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}decode_whole_run_tps 9.09 tok/s {"generation_seconds": 393.673, "generated_tokens": 3578.8}tasks_attempted 2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 1 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom no {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 578.9, "requests": 21, "tool_calls": 30, "completion_tokens": 3601, "occupied_max": 9732}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 21.1, "requests": 1, "tool_calls": 1, "completion_tokens": 0, "occupied_max": 0}] listtasks_over_64k [] instance idscompletion_tokens 3,601 tokens {"finish_reasons": {"tool_calls": 21}}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Agentic summary: 600 s, 30 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 8.85 tok/s, whole run 9.09 tok/s.
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit d4f093c724ebfd683feed823be35a11d4c009375 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T190629Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2-stability.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2
20261003T191656Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2-lab ! fail lab T20b eligibility: none failed stage: gates
Config llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 Started 2026-10-04 05:16 AEST Finished 2026-10-04 05:31 AEST (15.0 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T182602Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048.json sha256 89ee35a7fb8e6ec8fd63aa7e9e07a5ed1c3448eed3d10cebf357887e810a94feMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 9.50 tok/s lab_prefill_tps 199 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit d4f093c724ebfd683feed823be35a11d4c009375 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T191656Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2-lab.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2
20261003T193228Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2-load ✓ pass load T22b eligibility: none
Config ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2 Started 2026-10-04 05:32 AEST Finished 2026-10-04 05:33 AEST (0.9 min) Exit status None Notes –
Other runs of this config: load lab ! TTFT stability ✕
Launch recorded by tools/run.py (artifacts/runs/20261003T193228Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --ctx-checkpoints 4 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4.json sha256 aa948a723db410fdb552afee2e3ac6e4d2fa58a7a7d510ddd629de6bd166b13aMTP / speculative {"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --ctx-checkpoints 4env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 51.9 s host_ram_drop_gb 115 GB vram_ready_mib 22,318 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 956afbaae18f13c44809e1d61a0a1568dca9db78 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T193228Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2-load.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2
20261003T193321Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2-lab ! fail lab T22b eligibility: none failed stage: gates
Config ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2 Started 2026-10-04 05:33 AEST Finished 2026-10-04 05:48 AEST (15.0 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
Other runs of this config: load lab ! TTFT stability ✕
Launch recorded by tools/run.py (artifacts/runs/20261003T193228Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --ctx-checkpoints 4 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4.json sha256 aa948a723db410fdb552afee2e3ac6e4d2fa58a7a7d510ddd629de6bd166b13aMTP / speculative {"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --ctx-checkpoints 4env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": false, "context": true, "speed": false}lab_decode_c1_tps 8.60 tok/s lab_prefill_tps 147 tok/s "context-gate prompt tokens / total request seconds"mtp_acceptance_rate 0.5403 fraction {"draft_tokens": 2889, "accepted_tokens": 1561, "source": "server log delta"}mtp_accepted_tps 8.60 tok/s "output tok/s including accepted draft tokens"
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 956afbaae18f13c44809e1d61a0a1568dca9db78 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T193321Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2-lab.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2
20261003T194821Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2-ttft ✓ pass TTFT T22b eligibility: none
Config ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2 Started 2026-10-04 05:48 AEST Finished 2026-10-04 06:02 AEST (13.9 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab ! TTFT stability ✕
Launch recorded by tools/run.py (artifacts/runs/20261003T193228Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --ctx-checkpoints 4 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4.json sha256 aa948a723db410fdb552afee2e3ac6e4d2fa58a7a7d510ddd629de6bd166b13aMTP / speculative {"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --ctx-checkpoints 4env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 199 tok/s {"prompt_tokens": [10721, 10696, 10751], "ttft_s": [53.479, 53.836, 54.249]}ttft_prefill_32k_tps 191 tok/s {"prompt_tokens": [42813, 42832, 42764], "ttft_s": [223.74, 224.908, 224.251]}ttft_prefill_tps 199 tok/s {"prompt_tokens": [10721, 10696, 10751], "ttft_s": [53.479, 53.836, 54.249]}
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 956afbaae18f13c44809e1d61a0a1568dca9db78 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T194821Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2-ttft.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2
20261003T200216Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2-stability ✕ crash stability T22b eligibility: none failed stage: measure
Config ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2 Started 2026-10-04 06:02 AEST Finished 2026-10-04 06:12 AEST (10.1 min) Exit status 0 Notes stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: engine connection failures (crash/OOM); 9 tool-call parse failure(s); decode floor not verifiable (too little generation)
Other runs of this config: load lab ! TTFT stability ✕
Launch recorded by tools/run.py (artifacts/runs/20261003T193228Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --ctx-checkpoints 4 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4.json sha256 aa948a723db410fdb552afee2e3ac6e4d2fa58a7a7d510ddd629de6bd166b13aMTP / speculative {"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --ctx-checkpoints 4env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 1,641 tokens context.headroom_min 129,356 tokens duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 0 calls {"per_min": 0.0, "requests": 9}parse_failures 9 responses {"wire": {"unparsed_markup": 9}, "harness_format_errors": {"no_tool_call": 9}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 9 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}decode_window_min_tps unavailable only 58.214 s of generation; need > 120 s for one window decode_whole_run_tps 12.5 tok/s {"generation_seconds": 58.214, "generated_tokens": 729.8}tasks_attempted 5 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 0 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom yes {"upstream_connection_failures": 54, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 41.0, "requests": 3, "tool_calls": 0, "completion_tokens": 274, "occupied_max": 1559}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 36.8, "requests": 3, "tool_calls": 0, "completion_tokens": 248, "occupied_max": 1641}, {"attempt": 3, "round": 1, "instance_id": "psf__requests-2317", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 34.6, "requests": 3, "tool_calls": 0, "completion_tokens": 217, "occupied_max": 1609}, {"attempt": 4, "round": 1, "instance_id": "matplotlib__matplotlib-23299", "exit_status": "Error:BadGatewayError: litellm.BadGatewayError: BadGatewayError: OpenAIException - Cannot connect to host 127.0.0.1:18080 ssl:default [Connect call failed ('127.0.0.1', 18080)]", "submitted": false, "resolved": false, "seconds": 341.4}, {"attempt": 5, "round": 1, "instance_id": "scikit-learn__scikit-learn-13439", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 146.2}] listtasks_over_64k [] instance idscompletion_tokens 739 tokens {"finish_reasons": {"stop": 9}}mtp_accepted_tps unavailable not produced (crash/measure) mtp_acceptance_rate unavailable not produced (crash/measure)
Agentic summary: 600 s, 0 tool calls, 9 parse failures, 0 repetition hits, tasks solved 0/5, decode min window – tok/s, whole run 12.5 tok/s.
Cache, layout and host Cache warm Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-31 cpus_offline none machine_checks_this_boot 0
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 956afbaae18f13c44809e1d61a0a1568dca9db78 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T200216Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2-stability.json
Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012
20261003T220620Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012-load ✓ pass load T11 eligibility: none
Config llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012 Started 2026-10-04 08:06 AEST Finished 2026-10-04 08:07 AEST (0.7 min) Exit status None Notes –
Other runs of this config: load stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T220620Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit.json sha256 f7129eb274fa7a838d7a8928e9e72cec61e157259a237393b1f99b1945e722abMTP / speculative off argv /app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layerenv CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/Qwen3.8-Flash-Next-GGUF @928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 39.9 s host_ram_drop_gb 3.35 GB vram_ready_mib 68,976 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 2e517d99a59b71bbf630b44ce6204f3b85717adb cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261003T220620Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012-load.json
Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012
20261003T221725Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012-stability ! fail stability T11 eligibility: none
Config llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012 Started 2026-10-04 08:17 AEST Finished 2026-10-04 08:19 AEST (1.9 min) Exit status 0 Notes stability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode floor not verifiable (too little generation); agent driver exited 1
Other runs of this config: load stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T220620Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit.json sha256 f7129eb274fa7a838d7a8928e9e72cec61e157259a237393b1f99b1945e722abMTP / speculative off argv /app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layerenv CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/Qwen3.8-Flash-Next-GGUF @928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no successful request context.headroom_min unavailable configured context unknown (set CONFIGURED_CTX) duration_s unavailable driver produced no timing tool_calls 0 calls {"per_min": null, "requests": 0}parse_failures 0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 0 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": null, "saved": "artifacts: proxy/repetition/"}decode_window_min_tps unavailable only 0.0 s of generation; need > 120 s for one window decode_whole_run_tps unavailable no generation recorded tasks_attempted 0 attempts {"unfinished_at_cap": 0, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved skipped no attempts crash_or_oom no {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [] listtasks_over_64k [] instance idscompletion_tokens 0 tokens {"finish_reasons": {}}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Agentic summary: – s, 0 tool calls, 0 parse failures, 0 repetition hits, tasks solved –/0, decode min window – tok/s, whole run – tok/s.
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit d9944c1c5f811326cc98ee2701073db1981ab9b7 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261003T221725Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012-stability.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012
20261003T222524Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012-load ✓ pass load T21 eligibility: none
Config llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 Started 2026-10-04 08:25 AEST Finished 2026-10-04 08:27 AEST (1.7 min) Exit status None Notes –
Other runs of this config: load lab ! lab ✕ TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T222524Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048.json sha256 ada3d8f510a24f0d5811d7af83862f218e4475f6cd2935003b2162f2c40e76ceMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 101 s host_ram_drop_gb 40.4 GB vram_ready_mib 68,496 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit cf7ad504f20ebd9fc05fbafb689f4809d92ed210 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T222524Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012-load.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012
20261003T222706Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012-lab ! fail lab T21 eligibility: none failed stage: gates
Config llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 Started 2026-10-04 08:27 AEST Finished 2026-10-04 08:51 AEST (24.6 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
Other runs of this config: load lab ! lab ✕ TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T222524Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048.json sha256 ada3d8f510a24f0d5811d7af83862f218e4475f6cd2935003b2162f2c40e76ceMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 12.3 tok/s lab_prefill_tps 158 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit cf7ad504f20ebd9fc05fbafb689f4809d92ed210 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T222706Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012-lab.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012
20261003T225145Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012-ttft ✓ pass TTFT T21 eligibility: none
Config llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 Started 2026-10-04 08:51 AEST Finished 2026-10-04 09:07 AEST (15.5 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab ! lab ✕ TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T222524Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048.json sha256 ada3d8f510a24f0d5811d7af83862f218e4475f6cd2935003b2162f2c40e76ceMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 166 tok/s {"prompt_tokens": [10720, 10806, 10675], "ttft_s": [64.46, 64.694, 64.976]}ttft_prefill_32k_tps 175 tok/s {"prompt_tokens": [42882, 42861, 42885], "ttft_s": [244.773, 246.002, 245.262]}ttft_prefill_tps 166 tok/s {"prompt_tokens": [10720, 10806, 10675], "ttft_s": [64.46, 64.694, 64.976]}
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit cf7ad504f20ebd9fc05fbafb689f4809d92ed210 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T225145Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012-ttft.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012
20261003T230715Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012-stability ! fail stability T21 eligibility: none
Config llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 Started 2026-10-04 09:07 AEST Finished 2026-10-04 09:17 AEST (10.5 min) Exit status 0 Notes stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 26 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 11.388 < floor 15.0
Other runs of this config: load lab ! lab ✕ TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T222524Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048.json sha256 ada3d8f510a24f0d5811d7af83862f218e4475f6cd2935003b2162f2c40e76ceMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 8,118 tokens context.headroom_min 122,865 tokens duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 26 calls {"per_min": 2.6, "requests": 15}parse_failures 0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 0 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}decode_window_min_tps 11.4 tok/s window 1: 11.57 tok/s window 2: 11.47 tok/s window 3: 11.64 tok/s min 11.47 · max 11.64 tok/s {"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 143, "min_at_generation_s": 109.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}decode_whole_run_tps 11.7 tok/s {"generation_seconds": 262.271, "generated_tokens": 3060.1}tasks_attempted 2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 1 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom no {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 364.3, "requests": 12, "tool_calls": 20, "completion_tokens": 2740, "occupied_max": 8118}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 235.7, "requests": 4, "tool_calls": 6, "completion_tokens": 336, "occupied_max": 2784}] listtasks_over_64k [] instance idscompletion_tokens 3,076 tokens {"finish_reasons": {"tool_calls": 15}}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Agentic summary: 600 s, 26 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 11.4 tok/s, whole run 11.7 tok/s.
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit cf7ad504f20ebd9fc05fbafb689f4809d92ed210 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T230715Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012-stability.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012
20261003T231744Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012-lab ✕ crash lab T21 eligibility: none failed stage: harness
Config llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 Started 2026-10-04 09:17 AEST Finished 2026-10-04 09:29 AEST (11.3 min) Exit status 1 Notes –
Other runs of this config: load lab ! lab ✕ TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T222524Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048.json sha256 ada3d8f510a24f0d5811d7af83862f218e4475f6cd2935003b2162f2c40e76ceMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run gates_passed unavailable not produced (crash/harness) lab_decode_c1_tps unavailable not produced (crash/harness) lab_prefill_tps unavailable not produced (crash/harness) mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit cf7ad504f20ebd9fc05fbafb689f4809d92ed210 cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T231744Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012-lab.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012
20261003T232915Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012-load ✓ pass load T21 eligibility: none
Config llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 Started 2026-10-04 09:29 AEST Finished 2026-10-04 09:30 AEST (1.3 min) Exit status None Notes –
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T232915Z-lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ3_XXS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048.json sha256 4112d34e61cb2bb3861a59a7bf7cc1ce23cbba5e7f1b7cb3ce73f36dcfa2b7afMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ3_XXS.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 79.7 s host_ram_drop_gb 2.99 GB vram_ready_mib 67,148 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit d6382f80d3a94ab1a97311c8a27427f2b5c7f9cb cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T232915Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012-load.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012
20261003T233035Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012-lab ! fail lab T21 eligibility: none failed stage: gates
Config llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 Started 2026-10-04 09:30 AEST Finished 2026-10-04 10:12 AEST (42.0 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T232915Z-lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ3_XXS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048.json sha256 4112d34e61cb2bb3861a59a7bf7cc1ce23cbba5e7f1b7cb3ce73f36dcfa2b7afMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ3_XXS.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 10.5 tok/s lab_prefill_tps 126 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 94e598f979af99cb8e62f9dacf2a0f1b258ccf3e cleanpins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261003T233035Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012-lab.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012
20261004T001236Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012-ttft ✓ pass TTFT T21 eligibility: none
Config llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 Started 2026-10-04 10:12 AEST Finished 2026-10-04 10:31 AEST (19.3 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T232915Z-lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ3_XXS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048.json sha256 4112d34e61cb2bb3861a59a7bf7cc1ce23cbba5e7f1b7cb3ce73f36dcfa2b7afMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ3_XXS.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 130 tok/s {"prompt_tokens": [10633, 10740, 10676], "ttft_s": [80.845, 98.056, 82.319]}ttft_prefill_32k_tps 143 tok/s {"prompt_tokens": [42717, 42962, 42901], "ttft_s": [298.369, 301.101, 295.768]}ttft_prefill_tps 130 tok/s {"prompt_tokens": [10633, 10740, 10676], "ttft_s": [80.845, 98.056, 82.319]}
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 376b9064b928086c6007ed16ab0d9267e111dd92 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T001236Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012-ttft.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012
20261004T003153Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012-stability ! fail stability T21 eligibility: none
Config llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 Started 2026-10-04 10:31 AEST Finished 2026-10-04 10:42 AEST (10.5 min) Exit status 0 Notes stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 22 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: 1 repetition hit(s); decode window min 9.855 < floor 15.0
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T232915Z-lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ3_XXS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048.json sha256 4112d34e61cb2bb3861a59a7bf7cc1ce23cbba5e7f1b7cb3ce73f36dcfa2b7afMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ3_XXS.json
Metrics context.configured 131,072 tokens context.occupied_max 5,309 tokens context.headroom_min 125,652 tokens duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 22 calls {"per_min": 2.2, "requests": 16}parse_failures 0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 0 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 1 hits {"ngram64x3": 1, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}decode_window_min_tps 9.86 tok/s window 1: 10.15 tok/s window 2: 10.0 tok/s window 3: 10.01 tok/s min 10 · max 10.15 tok/s {"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 148, "min_at_generation_s": 195.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}decode_whole_run_tps 10.1 tok/s {"generation_seconds": 267.587, "generated_tokens": 2697.7}tasks_attempted 2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 1 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom no {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 225.2, "requests": 10, "tool_calls": 15, "completion_tokens": 1117, "occupied_max": 3970}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 374.8, "requests": 7, "tool_calls": 7, "completion_tokens": 1598, "occupied_max": 5309}] listtasks_over_64k [] instance idscompletion_tokens 2,715 tokens {"finish_reasons": {"tool_calls": 16}}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Agentic summary: 600 s, 22 tool calls, 0 parse failures, 1 repetition hits, tasks solved 1/2, decode min window 9.86 tok/s, whole run 10.1 tok/s.
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 376b9064b928086c6007ed16ab0d9267e111dd92 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T003153Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012-stability.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012
20261004T004223Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012-lab ! fail lab T21 eligibility: none failed stage: gates
Config llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 Started 2026-10-04 10:42 AEST Finished 2026-10-04 11:09 AEST (27.2 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261003T232915Z-lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ3_XXS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048.json sha256 4112d34e61cb2bb3861a59a7bf7cc1ce23cbba5e7f1b7cb3ce73f36dcfa2b7afMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ3_XXS.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 10.4 tok/s lab_prefill_tps 132 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 376b9064b928086c6007ed16ab0d9267e111dd92 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T004223Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012-lab.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012
20261004T013855Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012-load ✓ pass load T21 eligibility: none
Config llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012 Started 2026-10-04 11:38 AEST Finished 2026-10-04 11:39 AEST (0.9 min) Exit status None Notes –
Other runs of this config: load lab ! TTFT
Launch recorded by tools/run.py (artifacts/runs/20261004T013855Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048.json sha256 d46ba43c1169b3fd2e5d00ace10dc78cf7122918e1c259debec506f8a0766d22MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 55.0 s host_ram_drop_gb 3.83 GB vram_ready_mib 65,124 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 640a2b5606870f5adba5d1a5f8a557a075bbdecc dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T013855Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012-load.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012
20261004T013950Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012-lab ! fail lab T21 eligibility: none failed stage: gates
Config llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012 Started 2026-10-04 11:39 AEST Finished 2026-10-04 11:59 AEST (19.8 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
Other runs of this config: load lab ! TTFT
Launch recorded by tools/run.py (artifacts/runs/20261004T013855Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048.json sha256 d46ba43c1169b3fd2e5d00ace10dc78cf7122918e1c259debec506f8a0766d22MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 12.2 tok/s lab_prefill_tps 153 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 640a2b5606870f5adba5d1a5f8a557a075bbdecc dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T013950Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012-lab.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012
20261004T015940Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012-ttft ✓ pass TTFT T21 eligibility: none
Config llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012 Started 2026-10-04 11:59 AEST Finished 2026-10-04 12:15 AEST (16.0 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab ! TTFT
Launch recorded by tools/run.py (artifacts/runs/20261004T013855Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048.json sha256 d46ba43c1169b3fd2e5d00ace10dc78cf7122918e1c259debec506f8a0766d22MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 159 tok/s {"prompt_tokens": [10694, 10665, 10720], "ttft_s": [67.037, 66.91, 67.278]}ttft_prefill_32k_tps 169 tok/s {"prompt_tokens": [42746, 42816, 42870], "ttft_s": [254.159, 253.337, 253.851]}ttft_prefill_tps 159 tok/s {"prompt_tokens": [10694, 10665, 10720], "ttft_s": [67.037, 66.91, 67.278]}
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 640a2b5606870f5adba5d1a5f8a557a075bbdecc dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T015940Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012-ttft.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012
20261004T021546Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012-load ✓ pass load T21 eligibility: none
Config llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012 Started 2026-10-04 12:15 AEST Finished 2026-10-04 12:16 AEST (0.9 min) Exit status None Notes –
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T021546Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-g/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-g --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210.json sha256 1f7fcd2dbb10fca06661e2443bf4355bb609fd7677b4665a7891f97c2f975aa9MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 54.9 s host_ram_drop_gb 4.03 GB vram_ready_mib 65,124 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit fd89b522703b1628ea955d1581e800faae3e8e5f dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T021546Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012-load.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012
20261004T021642Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012-lab ! fail lab T21 eligibility: none failed stage: gates
Config llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012 Started 2026-10-04 12:16 AEST Finished 2026-10-04 12:27 AEST (11.2 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T021546Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-g/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-g --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210.json sha256 1f7fcd2dbb10fca06661e2443bf4355bb609fd7677b4665a7891f97c2f975aa9MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 11.8 tok/s lab_prefill_tps 219 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit fd89b522703b1628ea955d1581e800faae3e8e5f dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T021642Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012-lab.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012
20261004T022756Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012-ttft ✓ pass TTFT T21 eligibility: none
Config llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012 Started 2026-10-04 12:27 AEST Finished 2026-10-04 12:39 AEST (11.1 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T021546Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-g/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-g --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210.json sha256 1f7fcd2dbb10fca06661e2443bf4355bb609fd7677b4665a7891f97c2f975aa9MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 240 tok/s {"prompt_tokens": [10656, 10730, 10756], "ttft_s": [44.441, 45.48, 44.34]}ttft_prefill_32k_tps 242 tok/s {"prompt_tokens": [42911, 42857, 42784], "ttft_s": [177.67, 176.719, 177.974]}ttft_prefill_tps 240 tok/s {"prompt_tokens": [10656, 10730, 10756], "ttft_s": [44.441, 45.48, 44.34]}
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit fd89b522703b1628ea955d1581e800faae3e8e5f dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T022756Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012-ttft.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012
20261004T023904Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012-stability ! fail stability T21 eligibility: none
Config llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012 Started 2026-10-04 12:39 AEST Finished 2026-10-04 12:49 AEST (10.4 min) Exit status 0 Notes stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 24 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 11.035 < floor 15.0
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T021546Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-g/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-g --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210.json sha256 1f7fcd2dbb10fca06661e2443bf4355bb609fd7677b4665a7891f97c2f975aa9MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 6,174 tokens context.headroom_min 124,753 tokens duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 24 calls {"per_min": 2.4, "requests": 16}parse_failures 0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 0 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}decode_window_min_tps 11.0 tok/s window 1: 11.25 tok/s window 2: 11.07 tok/s min 11.07 · max 11.25 tok/s {"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 78, "min_at_generation_s": 109.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}decode_whole_run_tps 11.3 tok/s {"generation_seconds": 197.184, "generated_tokens": 2222.9}tasks_attempted 2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 1 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom no {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 266.1, "requests": 13, "tool_calls": 19, "completion_tokens": 2034, "occupied_max": 6174}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 333.9, "requests": 4, "tool_calls": 5, "completion_tokens": 206, "occupied_max": 3300}] listtasks_over_64k [] instance idscompletion_tokens 2,240 tokens {"finish_reasons": {"tool_calls": 16}}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Agentic summary: 600 s, 24 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 11.0 tok/s, whole run 11.3 tok/s.
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit fd89b522703b1628ea955d1581e800faae3e8e5f dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T023904Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012-stability.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012
20261004T024930Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012-lab ! fail lab T21 eligibility: none failed stage: gates
Config llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012 Started 2026-10-04 12:49 AEST Finished 2026-10-04 13:05 AEST (16.4 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T021546Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-g/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-g --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210.json sha256 1f7fcd2dbb10fca06661e2443bf4355bb609fd7677b4665a7891f97c2f975aa9MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 11.8 tok/s lab_prefill_tps 218 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit fd89b522703b1628ea955d1581e800faae3e8e5f dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T024930Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012-lab.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2
20261004T030555Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2-load ✓ pass load T20b eligibility: none
Config llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2 Started 2026-10-04 13:05 AEST Finished 2026-10-04 13:06 AEST (0.8 min) Exit status None Notes –
Other runs of this config: load lab ! TTFT
Launch recorded by tools/run.py (artifacts/runs/20261004T030555Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15.json sha256 72b9b0b1ecb00a8819e5a83e80deef181225276cbaa232ea9d7ed944c7391289MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 45.9 s host_ram_drop_gb 3.52 GB vram_ready_mib 21,944 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 0ace6784b9e48161d6b2ab80af2c1887ace48181 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T030555Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2-load.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2
20261004T030642Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2-lab ! fail lab T20b eligibility: none failed stage: gates
Config llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2 Started 2026-10-04 13:06 AEST Finished 2026-10-04 13:29 AEST (22.3 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
Other runs of this config: load lab ! TTFT
Launch recorded by tools/run.py (artifacts/runs/20261004T030555Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15.json sha256 72b9b0b1ecb00a8819e5a83e80deef181225276cbaa232ea9d7ed944c7391289MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 9.50 tok/s lab_prefill_tps 201 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit bc3d891210f23ffb6504b865b4216f82e08a42ef dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T030642Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2-lab.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2
20261004T032900Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2-ttft ✓ pass TTFT T20b eligibility: none
Config llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2 Started 2026-10-04 13:29 AEST Finished 2026-10-04 13:41 AEST (12.2 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab ! TTFT
Launch recorded by tools/run.py (artifacts/runs/20261004T030555Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15.json sha256 72b9b0b1ecb00a8819e5a83e80deef181225276cbaa232ea9d7ed944c7391289MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 213 tok/s {"prompt_tokens": [10652, 10738, 10639], "ttft_s": [50.031, 50.271, 51.458]}ttft_prefill_32k_tps 222 tok/s {"prompt_tokens": [42844, 42844, 42698], "ttft_s": [192.579, 195.095, 192.023]}ttft_prefill_tps 213 tok/s {"prompt_tokens": [10652, 10738, 10639], "ttft_s": [50.031, 50.271, 51.458]}
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit bc3d891210f23ffb6504b865b4216f82e08a42ef dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T032900Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2-ttft.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2
20261004T034114Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2-load ✓ pass load T20b eligibility: none
Config llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 Started 2026-10-04 13:41 AEST Finished 2026-10-04 13:42 AEST (0.8 min) Exit status None Notes –
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T034114Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 4096 --ubatch-size 4096 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096.json sha256 ee2f2a2da672b25768c3764ea94939d349c3e8c3616b4ebf5a18302b922972beMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 4096 --ubatch-size 4096env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 45.9 s host_ram_drop_gb 4.55 GB vram_ready_mib 22,914 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 2dbf31a09288732fa781855672b4d25178eb4682 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T034114Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2-load.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2
20261004T034201Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2-lab ! fail lab T20b eligibility: none failed stage: gates
Config llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 Started 2026-10-04 13:42 AEST Finished 2026-10-04 13:55 AEST (13.9 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T034114Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 4096 --ubatch-size 4096 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096.json sha256 ee2f2a2da672b25768c3764ea94939d349c3e8c3616b4ebf5a18302b922972beMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 4096 --ubatch-size 4096env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 9.60 tok/s lab_prefill_tps 288 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 2dbf31a09288732fa781855672b4d25178eb4682 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T034201Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2-lab.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2
20261004T035555Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2-ttft ✓ pass TTFT T20b eligibility: none
Config llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 Started 2026-10-04 13:55 AEST Finished 2026-10-04 14:04 AEST (8.5 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T034114Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 4096 --ubatch-size 4096 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096.json sha256 ee2f2a2da672b25768c3764ea94939d349c3e8c3616b4ebf5a18302b922972beMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 4096 --ubatch-size 4096env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 321 tok/s {"prompt_tokens": [10699, 10766, 10755], "ttft_s": [33.318, 33.541, 34.552]}ttft_prefill_32k_tps 316 tok/s {"prompt_tokens": [42816, 42866, 42738], "ttft_s": [135.262, 135.596, 135.444]}ttft_prefill_tps 321 tok/s {"prompt_tokens": [10699, 10766, 10755], "ttft_s": [33.318, 33.541, 34.552]}
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 2dbf31a09288732fa781855672b4d25178eb4682 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T035555Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2-ttft.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2
20261004T040423Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2-stability ! fail stability T20b eligibility: none
Config llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 Started 2026-10-04 14:04 AEST Finished 2026-10-04 14:14 AEST (10.5 min) Exit status 0 Notes stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 27 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: 1 repetition hit(s); decode window min 8.847 < floor 15.0
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T034114Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 4096 --ubatch-size 4096 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096.json sha256 ee2f2a2da672b25768c3764ea94939d349c3e8c3616b4ebf5a18302b922972beMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 4096 --ubatch-size 4096env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 6,632 tokens context.headroom_min 124,337 tokens duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 27 calls {"per_min": 2.7, "requests": 17}parse_failures 0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 0 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 1 hits {"ngram64x3": 1, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}decode_window_min_tps 8.85 tok/s window 1: 9.1 tok/s window 2: 8.89 tok/s window 3: 9.14 tok/s window 4: 9.3 tok/s window 5: 9.15 tok/s window 6: 9.01 tok/s min 8.89 · max 9.3 tok/s {"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 314, "min_at_generation_s": 123.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}decode_whole_run_tps 9.10 tok/s {"generation_seconds": 433.213, "generated_tokens": 3940.0}tasks_attempted 2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 1 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom no {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 331.4, "requests": 11, "tool_calls": 18, "completion_tokens": 2064, "occupied_max": 6088}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 268.6, "requests": 7, "tool_calls": 9, "completion_tokens": 1894, "occupied_max": 6632}] listtasks_over_64k [] instance idscompletion_tokens 3,958 tokens {"finish_reasons": {"tool_calls": 17}}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Agentic summary: 600 s, 27 tool calls, 0 parse failures, 1 repetition hits, tasks solved 1/2, decode min window 8.85 tok/s, whole run 9.10 tok/s.
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 2dbf31a09288732fa781855672b4d25178eb4682 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T040423Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2-stability.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2
20261004T041451Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2-lab ! fail lab T20b eligibility: none failed stage: gates
Config llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 Started 2026-10-04 14:14 AEST Finished 2026-10-04 14:23 AEST (8.6 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T034114Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 4096 --ubatch-size 4096 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096.json sha256 ee2f2a2da672b25768c3764ea94939d349c3e8c3616b4ebf5a18302b922972beMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 4096 --ubatch-size 4096env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 9.60 tok/s lab_prefill_tps 288 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 2dbf31a09288732fa781855672b4d25178eb4682 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T041451Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2-lab.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-gpu012
20261004T042329Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-gpu012-load ✕ crash load T22c eligibility: none failed stage: load
Config ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-gpu012 Started 2026-10-04 14:23 AEST Finished 2026-10-04 14:24 AEST (0.7 min) Exit status 1 Notes che_init: CUDA2 KV buffer size = 570.81 MiB
llama_init_from_model: KV self size = 748.00 MiB, c^KV (q8_0): 748.00 MiB, kv^T: not used
llama_init_from_model: KV self indexer size = 704.00 MiB (f16)
llama_init_from_model: CUDA_Host output buffer size = 0.59 MiB
ggml_backend_cuda_buffer_type_alloc_buffer: allocating 3488.51 MiB on device 0: cudaMalloc failed: out of memory
ggml_gallocr_reserve_n: failed to allocate CUDA0 buffer of size 3657965696
llama_init_from_model: failed to allocate compute buffers
~ggml_backend_cuda_context: have 0 graphs
~ggml_backend_cuda_context: have 0 graphs
~ggml_backend_cuda_context: have 0 graphs
llama_init_from_gpt_params: error: failed to create context with model '/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf'
Launch recorded by tools/run.py (artifacts/runs/20261004T042329Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-gpu012/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210.json sha256 b1d4a0440963d486728825e444f7ac6736f7d39444e483cdf843d738cf41f923MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds unavailable not produced (crash/load) host_ram_drop_gb unavailable not produced (crash/load) vram_ready_mib unavailable not produced (crash/load)
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 562ad23a6a0d69a48c21a7611e7c2d9f0984a1ed dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T042329Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-gpu012-load.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-gpu012
20261004T042414Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-gpu012-load ✕ crash load T22c eligibility: none failed stage: load
Config ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-gpu012 Started 2026-10-04 14:24 AEST Finished 2026-10-04 14:25 AEST (0.8 min) Exit status 1 Notes che_init: CUDA2 KV buffer size = 570.81 MiB
llama_init_from_model: KV self size = 816.00 MiB, c^KV (q8_0): 816.00 MiB, kv^T: not used
llama_init_from_model: KV self indexer size = 768.00 MiB (f16)
llama_init_from_model: CUDA_Host output buffer size = 0.61 MiB
ggml_backend_cuda_buffer_type_alloc_buffer: allocating 3488.51 MiB on device 0: cudaMalloc failed: out of memory
ggml_gallocr_reserve_n: failed to allocate CUDA0 buffer of size 3657965696
llama_init_from_model: failed to allocate compute buffers
~ggml_backend_cuda_context: have 0 graphs
~ggml_backend_cuda_context: have 0 graphs
~ggml_backend_cuda_context: have 0 graphs
llama_init_from_gpt_params: error: failed to create context with model '/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf'
Launch recorded by tools/run.py (artifacts/runs/20261004T042414Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-gpu012/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210.json sha256 9f2b351bbcb5a9f58f9c86180a6023d2ce919c2149f9640d3778c925a0ec8881MTP / speculative {"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds unavailable not produced (crash/load) host_ram_drop_gb unavailable not produced (crash/load) vram_ready_mib unavailable not produced (crash/load)
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 0c585c1d628c0c4a04e8994c51e5954b718fabea dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T042414Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-gpu012-load.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012
20261004T042504Z-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012-load ✓ pass load T21 eligibility: none
Config llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012 Started 2026-10-04 14:25 AEST Finished 2026-10-04 14:26 AEST (1.5 min) Exit status None Notes –
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T042504Z-lab-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210- --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210.json sha256 6a17b74e230df72d3cf6d97693f004f278ead699c76a664b7f92b5fe36813807MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 91.2 s host_ram_drop_gb 3.05 GB vram_ready_mib 63,084 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 743ba9ddf00e0b7eb28b9180241f4a3fb9322a06 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T042504Z-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012-load.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012
20261004T042635Z-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012-lab ! fail lab T21 eligibility: none failed stage: gates
Config llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012 Started 2026-10-04 14:26 AEST Finished 2026-10-04 14:52 AEST (25.7 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T042504Z-lab-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210- --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210.json sha256 6a17b74e230df72d3cf6d97693f004f278ead699c76a664b7f92b5fe36813807MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 8.00 tok/s lab_prefill_tps 139 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 743ba9ddf00e0b7eb28b9180241f4a3fb9322a06 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T042635Z-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012-lab.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012
20261004T045220Z-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012-ttft ✓ pass TTFT T21 eligibility: none
Config llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012 Started 2026-10-04 14:52 AEST Finished 2026-10-04 15:08 AEST (16.7 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T042504Z-lab-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210- --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210.json sha256 6a17b74e230df72d3cf6d97693f004f278ead699c76a664b7f92b5fe36813807MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 149 tok/s {"prompt_tokens": [10777, 10733, 10753], "ttft_s": [130.896, 71.944, 64.716]}ttft_prefill_32k_tps 176 tok/s {"prompt_tokens": [42666, 42838, 42755], "ttft_s": [243.017, 245.943, 242.684]}ttft_prefill_tps 149 tok/s {"prompt_tokens": [10777, 10733, 10753], "ttft_s": [130.896, 71.944, 64.716]}
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 743ba9ddf00e0b7eb28b9180241f4a3fb9322a06 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T045220Z-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012-ttft.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012
20261004T050859Z-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012-stability ! fail stability T21 eligibility: none
Config llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012 Started 2026-10-04 15:08 AEST Finished 2026-10-04 15:19 AEST (10.5 min) Exit status 0 Notes stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 25 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 7.807 < floor 15.0
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T042504Z-lab-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210- --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210.json sha256 6a17b74e230df72d3cf6d97693f004f278ead699c76a664b7f92b5fe36813807MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max 6,852 tokens context.headroom_min 123,939 tokens duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 25 calls {"per_min": 2.5, "requests": 19}parse_failures 0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 0 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}decode_window_min_tps 7.81 tok/s window 1: 7.91 tok/s window 2: 7.96 tok/s window 3: 8.06 tok/s min 7.91 · max 8.06 tok/s {"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 175, "min_at_generation_s": 107.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}decode_whole_run_tps 7.95 tok/s {"generation_seconds": 294.515, "generated_tokens": 2340.5}tasks_attempted 2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 1 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom no {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 277.5, "requests": 9, "tool_calls": 12, "completion_tokens": 1323, "occupied_max": 4597}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 322.5, "requests": 11, "tool_calls": 13, "completion_tokens": 1038, "occupied_max": 6852}] listtasks_over_64k [] instance idscompletion_tokens 2,361 tokens {"finish_reasons": {"tool_calls": 19}}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Agentic summary: 600 s, 25 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 7.81 tok/s, whole run 7.95 tok/s.
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 743ba9ddf00e0b7eb28b9180241f4a3fb9322a06 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T050859Z-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012-stability.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012
20261004T051928Z-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012-lab ! fail lab T21 eligibility: none failed stage: gates
Config llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012 Started 2026-10-04 15:19 AEST Finished 2026-10-04 15:38 AEST (19.0 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T042504Z-lab-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210- --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210.json sha256 6a17b74e230df72d3cf6d97693f004f278ead699c76a664b7f92b5fe36813807MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 8.30 tok/s lab_prefill_tps 156 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 743ba9ddf00e0b7eb28b9180241f4a3fb9322a06 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T051928Z-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012-lab.json
Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012
20261004T053844Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012-load ✓ pass load T11 eligibility: none
Config llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012 Started 2026-10-04 15:38 AEST Finished 2026-10-04 15:39 AEST (0.7 min) Exit status None Notes –
Other runs of this config: load lab lab TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T053844Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210.json sha256 75dbaf82a0bd6bfa005b5ab1a0c00ed1eddb44fb5fe1060d7cbbd834d962f115MTP / speculative off argv /app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/Qwen3.8-Flash-Next-GGUF @928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 39.9 s host_ram_drop_gb 2.65 GB vram_ready_mib 65,510 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit fd8fc15971beec0df168ae2ee8c6c01aae2a7453 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T053844Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012-load.json
Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012
20261004T053925Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012-lab ✓ pass lab T11 eligibility: none
Config llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012 Started 2026-10-04 15:39 AEST Finished 2026-10-04 15:45 AEST (6.3 min) Exit status 0 Notes served=/models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf
Other runs of this config: load lab lab TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T053844Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210.json sha256 75dbaf82a0bd6bfa005b5ab1a0c00ed1eddb44fb5fe1060d7cbbd834d962f115MTP / speculative off argv /app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/Qwen3.8-Flash-Next-GGUF @928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max 106,294 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}lab_decode_c1_tps 56.9 tok/s lab_prefill_tps 605 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit e17f05543581ac460774a08531a895566f2d209d dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T053925Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012-lab.json
Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012
20261004T054545Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012-ttft ✓ pass TTFT T11 eligibility: none
Config llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012 Started 2026-10-04 15:45 AEST Finished 2026-10-04 15:49 AEST (3.9 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab lab TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T053844Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210.json sha256 75dbaf82a0bd6bfa005b5ab1a0c00ed1eddb44fb5fe1060d7cbbd834d962f115MTP / speculative off argv /app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/Qwen3.8-Flash-Next-GGUF @928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 987 tok/s {"prompt_tokens": [13709, 13794, 13718], "ttft_s": [14.16, 13.87, 13.896]}ttft_prefill_32k_tps 852 tok/s {"prompt_tokens": [54701, 54603, 54830], "ttft_s": [64.015, 64.121, 64.836]}ttft_prefill_tps 987 tok/s {"prompt_tokens": [13709, 13794, 13718], "ttft_s": [14.16, 13.87, 13.896]}
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit e17f05543581ac460774a08531a895566f2d209d dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T054545Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012-ttft.json
Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012
20261004T054940Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012-stability ! fail stability T11 eligibility: none
Config llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012 Started 2026-10-04 15:49 AEST Finished 2026-10-04 16:00 AEST (10.4 min) Exit status 0 Notes stability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 70 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 38.626 < floor 50.0
Other runs of this config: load lab lab TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T053844Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210.json sha256 75dbaf82a0bd6bfa005b5ab1a0c00ed1eddb44fb5fe1060d7cbbd834d962f115MTP / speculative off argv /app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/Qwen3.8-Flash-Next-GGUF @928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max 39,226 tokens context.headroom_min 91,731 tokens duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 70 calls {"per_min": 7.0, "requests": 54}parse_failures 0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 0 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}decode_window_min_tps 38.6 tok/s window 1: 48.84 tok/s window 2: 45.07 tok/s window 3: 41.96 tok/s window 4: 39.7 tok/s min 39.7 · max 48.84 tok/s {"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 209, "min_at_generation_s": 267.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}decode_whole_run_tps 45.4 tok/s {"generation_seconds": 328.55, "generated_tokens": 14907.4}tasks_attempted 2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 1 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom no {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 49.5, "requests": 11, "tool_calls": 16, "completion_tokens": 2006, "occupied_max": 5874}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 550.5, "requests": 44, "tool_calls": 54, "completion_tokens": 12968, "occupied_max": 39226}] listtasks_over_64k [] instance idscompletion_tokens 14,974 tokens {"finish_reasons": {"tool_calls": 54}}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Agentic summary: 600 s, 70 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 38.6 tok/s, whole run 45.4 tok/s.
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit e17f05543581ac460774a08531a895566f2d209d dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T054940Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012-stability.json
Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012
20261004T060007Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012-lab ✓ pass lab T11 eligibility: none
Config llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012 Started 2026-10-04 16:00 AEST Finished 2026-10-04 16:13 AEST (13.3 min) Exit status 0 Notes served=/models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf
Other runs of this config: load lab lab TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T053844Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210.json sha256 75dbaf82a0bd6bfa005b5ab1a0c00ed1eddb44fb5fe1060d7cbbd834d962f115MTP / speculative off argv /app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/Qwen3.8-Flash-Next-GGUF @928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max 106,294 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}lab_decode_c1_tps 57.1 tok/s lab_prefill_tps 608 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit e17f05543581ac460774a08531a895566f2d209d dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T060007Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012-lab.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2
20261004T061329Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2-load ✓ pass load T20b eligibility: none
Config llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2 Started 2026-10-04 16:13 AEST Finished 2026-10-04 16:14 AEST (0.8 min) Exit status None Notes –
Other runs of this config: load lab ! TTFT
Launch recorded by tools/run.py (artifacts/runs/20261004T061329Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 8192 --ubatch-size 8192 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192.json sha256 6a9c08474b62518d1e1c1d23f6b6b1f241e7a0e05c5faa5f176e3700eec9fefeMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 8192 --ubatch-size 8192env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 48.9 s host_ram_drop_gb 5.84 GB vram_ready_mib 22,186 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 5af5149471babca3e64d893eeea35a3a1a43fe06 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T061329Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2-load.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2
20261004T061418Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2-lab ! fail lab T20b eligibility: none failed stage: gates
Config llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2 Started 2026-10-04 16:14 AEST Finished 2026-10-04 16:23 AEST (9.0 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
Other runs of this config: load lab ! TTFT
Launch recorded by tools/run.py (artifacts/runs/20261004T061329Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 8192 --ubatch-size 8192 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192.json sha256 6a9c08474b62518d1e1c1d23f6b6b1f241e7a0e05c5faa5f176e3700eec9fefeMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 8192 --ubatch-size 8192env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 9.10 tok/s lab_prefill_tps 370 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit dd520bfd96d949a57962b56ed6cd8dd55f834916 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T061418Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2-lab.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2
20261004T062318Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2-ttft ✓ pass TTFT T20b eligibility: none
Config llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2 Started 2026-10-04 16:23 AEST Finished 2026-10-04 16:30 AEST (6.7 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab ! TTFT
Launch recorded by tools/run.py (artifacts/runs/20261004T061329Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 8192 --ubatch-size 8192 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192.json sha256 6a9c08474b62518d1e1c1d23f6b6b1f241e7a0e05c5faa5f176e3700eec9fefeMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 8192 --ubatch-size 8192env CUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 384 tok/s {"prompt_tokens": [10746, 10734, 10753], "ttft_s": [27.965, 27.98, 27.965]}ttft_prefill_32k_tps 404 tok/s {"prompt_tokens": [42688, 42881, 42819], "ttft_s": [105.798, 105.946, 106.334]}ttft_prefill_tps 384 tok/s {"prompt_tokens": [10746, 10734, 10753], "ttft_s": [27.965, 27.98, 27.965]}
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit dd520bfd96d949a57962b56ed6cd8dd55f834916 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T062318Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2-ttft.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012
20261004T063003Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012-load ✓ pass load T21 eligibility: none
Config llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012 Started 2026-10-04 16:30 AEST Finished 2026-10-04 16:30 AEST (0.9 min) Exit status None Notes –
Other runs of this config: load lab ! TTFT
Launch recorded by tools/run.py (artifacts/runs/20261004T063003Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-g/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-g --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 4096 --ubatch-size 4096 --fit-target 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210.json sha256 e0b729222e80461545b33c6e10b8ff1d8ecaaf2fde1596824ce2a3e98b0ce76cMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 4096 --ubatch-size 4096 --fit-target 2048env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 54.9 s host_ram_drop_gb 4.93 GB vram_ready_mib 65,688 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 9a178e2fe77cc647e9542e1742a52c38931f07e1 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T063003Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012-load.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012
20261004T063058Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012-lab ! fail lab T21 eligibility: none failed stage: gates
Config llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012 Started 2026-10-04 16:30 AEST Finished 2026-10-04 16:48 AEST (17.9 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
Other runs of this config: load lab ! TTFT
Launch recorded by tools/run.py (artifacts/runs/20261004T063003Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-g/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-g --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 4096 --ubatch-size 4096 --fit-target 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210.json sha256 e0b729222e80461545b33c6e10b8ff1d8ecaaf2fde1596824ce2a3e98b0ce76cMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 4096 --ubatch-size 4096 --fit-target 2048env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 11.2 tok/s lab_prefill_tps 258 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 9a178e2fe77cc647e9542e1742a52c38931f07e1 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T063058Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012-lab.json
Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012
20261004T064850Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012-ttft ✓ pass TTFT T21 eligibility: none
Config llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012 Started 2026-10-04 16:48 AEST Finished 2026-10-04 16:58 AEST (9.3 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab ! TTFT
Launch recorded by tools/run.py (artifacts/runs/20261004T063003Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-g/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-g --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 4096 --ubatch-size 4096 --fit-target 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210.json sha256 e0b729222e80461545b33c6e10b8ff1d8ecaaf2fde1596824ce2a3e98b0ce76cMTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 4096 --ubatch-size 4096 --fit-target 2048env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 298 tok/s {"prompt_tokens": [10829, 10677, 10700], "ttft_s": [37.303, 35.77, 35.957]}ttft_prefill_32k_tps 288 tok/s {"prompt_tokens": [42788, 42675, 42983], "ttft_s": [148.453, 148.434, 149.751]}ttft_prefill_tps 298 tok/s {"prompt_tokens": [10829, 10677, 10700], "ttft_s": [37.303, 35.77, 35.957]}
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 9a178e2fe77cc647e9542e1742a52c38931f07e1 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T064850Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012-ttft.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012
20261004T065809Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012-load ✓ pass load T22c eligibility: none
Config ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 Started 2026-10-04 16:58 AEST Finished 2026-10-04 16:58 AEST (0.7 min) Exit status None Notes –
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T065809Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --fit-margin 4096 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096.json sha256 fe1c037653b578219259df5ac2f7fce9deed7442d4759cd749f0001ffe3df341MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --fit-margin 4096env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 44.3 s host_ram_drop_gb 65.5 GB vram_ready_mib 68,156 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 981b077ba6f48c9b94c9bc0a253f95413b94b11e dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T065809Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012-load.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012
20261004T065853Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012-lab ! fail lab T22c eligibility: none failed stage: gates
Config ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 Started 2026-10-04 16:58 AEST Finished 2026-10-04 17:10 AEST (12.0 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T065809Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --fit-margin 4096 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096.json sha256 fe1c037653b578219259df5ac2f7fce9deed7442d4759cd749f0001ffe3df341MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --fit-margin 4096env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": false, "context": true, "speed": false}lab_decode_c1_tps 14.0 tok/s lab_prefill_tps 185 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 981b077ba6f48c9b94c9bc0a253f95413b94b11e dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T065853Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012-lab.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012
20261004T071053Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012-ttft ✓ pass TTFT T22c eligibility: none
Config ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 Started 2026-10-04 17:10 AEST Finished 2026-10-04 17:21 AEST (10.4 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T065809Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --fit-margin 4096 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096.json sha256 fe1c037653b578219259df5ac2f7fce9deed7442d4759cd749f0001ffe3df341MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --fit-margin 4096env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 283 tok/s {"prompt_tokens": [10728, 10711, 10734], "ttft_s": [37.814, 37.788, 37.935]}ttft_prefill_32k_tps 253 tok/s {"prompt_tokens": [42827, 42804, 42935], "ttft_s": [167.744, 169.791, 169.931]}ttft_prefill_tps 283 tok/s {"prompt_tokens": [10728, 10711, 10734], "ttft_s": [37.814, 37.788, 37.935]}
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 981b077ba6f48c9b94c9bc0a253f95413b94b11e dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T071053Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012-ttft.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012
20261004T072115Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012-stability ! fail stability T22c eligibility: none
Config ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 Started 2026-10-04 17:21 AEST Finished 2026-10-04 17:31 AEST (10.1 min) Exit status 0 Notes stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: 68 tool-call parse failure(s); decode window min 13.255 < floor 15.0
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T065809Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --fit-margin 4096 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096.json sha256 fe1c037653b578219259df5ac2f7fce9deed7442d4759cd749f0001ffe3df341MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --fit-margin 4096env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 1,874 tokens context.headroom_min 129,132 tokens duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 0 calls {"per_min": 0.0, "requests": 68}parse_failures 68 responses {"wire": {"unparsed_markup": 68}, "harness_format_errors": {"no_tool_call": 68}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 68 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}decode_window_min_tps 13.3 tok/s window 1: 13.31 tok/s window 2: 13.29 tok/s window 3: 13.29 tok/s window 4: 13.35 tok/s window 5: 13.34 tok/s min 13.29 · max 13.35 tok/s {"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 299, "min_at_generation_s": 112.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}decode_whole_run_tps 13.3 tok/s {"generation_seconds": 418.663, "generated_tokens": 5576.2}tasks_attempted 23 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 0 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom no {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 30.9, "requests": 3, "tool_calls": 0, "completion_tokens": 246, "occupied_max": 1559}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 27.9, "requests": 3, "tool_calls": 0, "completion_tokens": 217, "occupied_max": 1641}, {"attempt": 3, "round": 1, "instance_id": "psf__requests-2317", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 37.7, "requests": 3, "tool_calls": 0, "completion_tokens": 348, "occupied_max": 1609}, {"attempt": 4, "round": 1, "instance_id": "matplotlib__matplotlib-23299", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 31.6, "requests": 3, "tool_calls": 0, "completion_tokens": 263, "occupied_max": 1874}, {"attempt": 5, "round": 1, "instance_id": "scikit-learn__scikit-learn-13439", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 31.2, "requests": 3, "tool_calls": 0, "completion_tokens": 260, "occupied_max": 1770}, {"attempt": 6, "round": 1, "instance_id": "astropy__astropy-12907", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 25.1, "requests": 3, "tool_calls": 0, "completion_tokens": 181, "occupied_max": 1716}, {"attempt": 7, "round": 1, "instance_id": "pytest-dev__pytest-7373", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 27.6, "requests": 3, "tool_calls": 0, "completion_tokens": 214, "occupied_max": 1641}, {"attempt": 8, "round": 1, "instance_id": "sympy__sympy-20590", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 39.7, "requests": 3, "tool_calls": 0, "completion_tokens": 375, "occupied_max": 1578}, {"attempt": 9, "round": 2, "instance_id": "django__django-11099", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 23.2, "requests": 3, "tool_calls": 0, "completion_tokens": 240, "occupied_max": 1559}, {"attempt": 10, "round": 2, "instance_id": "sympy__sympy-21612", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 21.1, "requests": 3, "tool_calls": 0, "completion_tokens": 212, "occupied_max": 1641}, {"attempt": 11, "round": 2, "instance_id": "psf__requests-2317", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 26.7, "requests": 3, "tool_calls": 0, "completion_tokens": 286, "occupied_max": 1609}, {"attempt": 12, "round": 2, "instance_id": "matplotlib__matplotlib-23299", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 25.0, "requests": 3, "tool_calls": 0, "completion_tokens": 262, "occupied_max": 1874}, {"attempt": 13, "round": 2, "instance_id": "scikit-learn__scikit-learn-13439", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 24.6, "requests": 3, "tool_calls": 0, "completion_tokens": 257, "occupied_max": 1770}, {"attempt": 14, "round": 2, "instance_id": "astropy__astropy-12907", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 19.0, "requests": 3, "tool_calls": 0, "completion_tokens": 184, "occupied_max": 1716}, {"attempt": 15, "round": 2, "instance_id": "pytest-dev__pytest-7373", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 20.8, "requests": 3, "tool_calls": 0, "completion_tokens": 208, "occupied_max": 1641}, {"attempt": 16, "round": 2, "instance_id": "sympy__sympy-20590", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 29.2, "requests": 3, "tool_calls": 0, "completion_tokens": 319, "occupied_max": 1578}, {"attempt": 17, "round": 3, "instance_id": "django__django-11099", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 23.2, "requests": 3, "tool_calls": 0, "completion_tokens": 240, "occupied_max": 1559}, {"attempt": 18, "round": 3, "instance_id": "sympy__sympy-21612", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 21.1, "requests": 3, "tool_calls": 0, "completion_tokens": 212, "occupied_max": 1641}, {"attempt": 19, "round": 3, "instance_id": "psf__requests-2317", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 26.8, "requests": 3, "tool_calls": 0, "completion_tokens": 286, "occupied_max": 1609}, {"attempt": 20, "round": 3, "instance_id": "matplotlib__matplotlib-23299", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 25.0, "requests": 3, "tool_calls": 0, "completion_tokens": 262, "occupied_max": 1874}, {"attempt": 21, "round": 3, "instance_id": "scikit-learn__scikit-learn-13439", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 24.6, "requests": 3, "tool_calls": 0, "completion_tokens": 257, "occupied_max": 1770}, {"attempt": 22, "round": 3, "instance_id": "astropy__astropy-12907", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 19.0, "requests": 3, "tool_calls": 0, "completion_tokens": 184, "occupied_max": 1716}, {"attempt": 23, "round": 3, "instance_id": "pytest-dev__pytest-7373", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 19.0, "requests": 3, "tool_calls": 0, "completion_tokens": 133, "occupied_max": 1522}] listtasks_over_64k [] instance idscompletion_tokens 5,646 tokens {"finish_reasons": {"stop": 68}}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Agentic summary: 600 s, 0 tool calls, 68 parse failures, 0 repetition hits, tasks solved 0/23, decode min window 13.3 tok/s, whole run 13.3 tok/s.
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 981b077ba6f48c9b94c9bc0a253f95413b94b11e dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T072115Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012-stability.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012
20261004T073119Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012-lab ! fail lab T22c eligibility: none failed stage: gates
Config ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 Started 2026-10-04 17:31 AEST Finished 2026-10-04 17:42 AEST (11.2 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
Other runs of this config: load lab ! lab ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T065809Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --fit-margin 4096 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096.json sha256 fe1c037653b578219259df5ac2f7fce9deed7442d4759cd749f0001ffe3df341MTP / speculative off argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --fit-margin 4096env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": false, "context": true, "speed": false}lab_decode_c1_tps 13.4 tok/s lab_prefill_tps 189 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 981b077ba6f48c9b94c9bc0a253f95413b94b11e dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T073119Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012-lab.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm4096-gpu012
20261004T074233Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm4096-gpu012-load ✕ crash load T22c eligibility: none failed stage: load
Config ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm4096-gpu012 Started 2026-10-04 17:42 AEST Finished 2026-10-04 17:43 AEST (0.9 min) Exit status 1 Notes c^KV (q8_0): 68.00 MiB, kv^T: not used
llama_init_from_model: KV self indexer size = 64.00 MiB (f16)
llama_init_from_model: CUDA_Host output buffer size = 0.61 MiB
ggml_backend_cuda_buffer_type_alloc_buffer: allocating 2754.25 MiB on device 0: cudaMalloc failed: out of memory
ggml_gallocr_reserve_n: failed to allocate CUDA0 buffer of size 2888039424
llama_init_from_model: failed to allocate compute buffers
~ggml_backend_cuda_context: have 0 graphs
~ggml_backend_cuda_context: have 0 graphs
~ggml_backend_cuda_context: have 0 graphs
common_speculative_init: failed to create MTP context
srv init: failed to initialize recurrent speculative context
~ggml_backend_cuda_context: have 13 graphs
~ggml_backend_cuda_context: have 14 graphs
~ggml_backend_cuda_context: have 11 graphs
Launch recorded by tools/run.py (artifacts/runs/20261004T074233Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm4096-gpu012/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm4096-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --fit-margin 4096 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm4096.json sha256 d21078e38ec9e96da581e9003a2aaae041591909413dbf70738357df0829341eMTP / speculative {"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --fit-margin 4096env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds unavailable not produced (crash/load) host_ram_drop_gb unavailable not produced (crash/load) vram_ready_mib unavailable not produced (crash/load)
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit f3cbbe1b98ba109e9c4abaac8d34bc7b6e2775ae dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T074233Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm4096-gpu012-load.json
Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012
20261004T080048Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012-load ✓ pass load T11 eligibility: none
Config llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012 Started 2026-10-04 18:00 AEST Finished 2026-10-04 18:01 AEST (0.7 min) Exit status None Notes –
Other runs of this config: load lab TTFT
Launch recorded by tools/run.py (artifacts/runs/20261004T080048Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 2048 --ubatch-size 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048.json sha256 e882737bdeeadb764f93582d803f25626630b8b0a09fb20716d71d1c2c7be516MTP / speculative off argv /app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 2048 --ubatch-size 2048env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/Qwen3.8-Flash-Next-GGUF @928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 39.9 s host_ram_drop_gb 3.74 GB vram_ready_mib 65,396 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 65a204856ba6e28ca395fa6aab47a88e41c62a1a dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T080048Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012-load.json
Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012
20261004T080128Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012-lab ✓ pass lab T11 eligibility: none
Config llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012 Started 2026-10-04 18:01 AEST Finished 2026-10-04 18:06 AEST (5.5 min) Exit status 0 Notes served=/models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf
Other runs of this config: load lab TTFT
Launch recorded by tools/run.py (artifacts/runs/20261004T080048Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 2048 --ubatch-size 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048.json sha256 e882737bdeeadb764f93582d803f25626630b8b0a09fb20716d71d1c2c7be516MTP / speculative off argv /app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 2048 --ubatch-size 2048env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/Qwen3.8-Flash-Next-GGUF @928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max 106,294 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}lab_decode_c1_tps 44.3 tok/s lab_prefill_tps 717 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 65a204856ba6e28ca395fa6aab47a88e41c62a1a dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T080128Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012-lab.json
Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012
20261004T080657Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012-ttft ✓ pass TTFT T11 eligibility: none
Config llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012 Started 2026-10-04 18:06 AEST Finished 2026-10-04 18:10 AEST (3.6 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab TTFT
Launch recorded by tools/run.py (artifacts/runs/20261004T080048Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 2048 --ubatch-size 2048 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048.json sha256 e882737bdeeadb764f93582d803f25626630b8b0a09fb20716d71d1c2c7be516MTP / speculative off argv /app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 2048 --ubatch-size 2048env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/Qwen3.8-Flash-Next-GGUF @928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 1,067 tok/s {"prompt_tokens": [13695, 13734, 13847], "ttft_s": [12.963, 12.872, 12.965]}ttft_prefill_32k_tps 931 tok/s {"prompt_tokens": [54728, 54685, 54604], "ttft_s": [58.434, 58.71, 58.906]}ttft_prefill_tps 1,067 tok/s {"prompt_tokens": [13695, 13734, 13847], "ttft_s": [12.963, 12.872, 12.965]}
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 65a204856ba6e28ca395fa6aab47a88e41c62a1a dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T080657Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012-ttft.json
Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012
20261004T081034Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012-load ✓ pass load T11 eligibility: none
Config llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012 Started 2026-10-04 18:10 AEST Finished 2026-10-04 18:11 AEST (0.7 min) Exit status None Notes –
Other runs of this config: load lab TTFT
Launch recorded by tools/run.py (artifacts/runs/20261004T081034Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 4096 --ubatch-size 4096 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096.json sha256 bb5c0e11a6559307e48f7f8f285d3439f9c7872614fa95ee033b2d48ca370dbdMTP / speculative off argv /app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 4096 --ubatch-size 4096env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/Qwen3.8-Flash-Next-GGUF @928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 39.9 s host_ram_drop_gb 4.34 GB vram_ready_mib 65,508 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit bfcf5657ea4ed56f452dc532e3311cf77750c3d9 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T081034Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012-load.json
Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012
20261004T081114Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012-lab ✓ pass lab T11 eligibility: none
Config llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012 Started 2026-10-04 18:11 AEST Finished 2026-10-04 18:19 AEST (8.1 min) Exit status 0 Notes served=/models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf
Other runs of this config: load lab TTFT
Launch recorded by tools/run.py (artifacts/runs/20261004T081034Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 4096 --ubatch-size 4096 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096.json sha256 bb5c0e11a6559307e48f7f8f285d3439f9c7872614fa95ee033b2d48ca370dbdMTP / speculative off argv /app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 4096 --ubatch-size 4096env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/Qwen3.8-Flash-Next-GGUF @928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max 106,294 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}lab_decode_c1_tps 35.0 tok/s lab_prefill_tps 713 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit bfcf5657ea4ed56f452dc532e3311cf77750c3d9 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T081114Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012-lab.json
Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012
20261004T081919Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012-ttft ✓ pass TTFT T11 eligibility: none
Config llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012 Started 2026-10-04 18:19 AEST Finished 2026-10-04 18:23 AEST (3.8 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab TTFT
Launch recorded by tools/run.py (artifacts/runs/20261004T081034Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210/docker-run.txt) Copy
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 4096 --ubatch-size 4096 Image ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Digest sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5Engine llama.cpp Engine source – Launch file launches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096.json sha256 bb5c0e11a6559307e48f7f8f285d3439f9c7872614fa95ee033b2d48ca370dbdMTP / speculative off argv /app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 4096 --ubatch-size 4096env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/Qwen3.8-Flash-Next-GGUF @928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 980 tok/s {"prompt_tokens": [13680, 13707, 13810], "ttft_s": [14.13, 13.965, 14.096]}ttft_prefill_32k_tps 881 tok/s {"prompt_tokens": [54557, 54627, 54737], "ttft_s": [61.779, 62.069, 62.105]}ttft_prefill_tps 980 tok/s {"prompt_tokens": [13680, 13707, 13810], "ttft_s": [14.13, 13.965, 14.096]}
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit bfcf5657ea4ed56f452dc532e3311cf77750c3d9 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T081919Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012-ttft.json
Runs / none
20261004T091142Z-hw-bw-gpu1-current-hw_bw ✓ pass hw bw T05a eligibility: none
Config hw-bw-gpu1-current Started 2026-10-04 19:11 AEST Finished 2026-10-04 19:11 AEST (0.2 min) Exit status 0 Notes 1 GiB pinned transfers; full table in detail
Launch no container: static measurement
Metrics context.configured – context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run h2d_gbps 13.5 GB/s {"1MB_pin": 13.14, "1MB_page": 13.01, "4MB_pin": 13.39, "4MB_page": 13.12, "16MB_pin": 13.45, "16MB_page": 13.18, "64MB_pin": 13.47, "64MB_page": 13.18, "256MB_pin": 13.47, "256MB_page": 13.07, "1024MB_pin": 13.47, "1024MB_page": 13.12}d2h_gbps 13.2 GB/s {"1MB_pin": 12.95, "1MB_page": 6.3, "4MB_pin": 13.14, "4MB_page": 10.05, "16MB_pin": 13.2, "16MB_page": 11.16, "64MB_pin": 13.21, "64MB_page": 11.79, "256MB_pin": 13.21, "256MB_page": 12.04, "1024MB_pin": 13.22, "1024MB_page": 12.05}host_memcpy_gbps 16.6 GB/s
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 00709a719d4c6057fbeb8f52eadd20c2bbe2f46d dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/none/20261004T091142Z-hw-bw-gpu1-current-hw_bw.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012
20261004T091154Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012-load ✓ pass load T22c eligibility: none
Config ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012 Started 2026-10-04 19:11 AEST Finished 2026-10-04 19:12 AEST (0.8 min) Exit status None Notes –
Other runs of this config: load lab ! TTFT stability ✕
Launch recorded by tools/run.py (artifacts/runs/20261004T091154Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --fit-margin 7168 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168.json sha256 cf6116ffb2e3698333ab48e1435f99a6f037b39308f84e62e4fd24888017f06cMTP / speculative {"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --fit-margin 7168env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 48.6 s host_ram_drop_gb 83.3 GB vram_ready_mib 64,680 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit c36c7e58741670e5e869f6857478b1e198932e9a dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T091154Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012-load.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012
20261004T091243Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012-lab ! fail lab T22c eligibility: none failed stage: gates
Config ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012 Started 2026-10-04 19:12 AEST Finished 2026-10-04 19:26 AEST (13.5 min) Exit status 0 Notes served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
Other runs of this config: load lab ! TTFT stability ✕
Launch recorded by tools/run.py (artifacts/runs/20261004T091154Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --fit-margin 7168 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168.json sha256 cf6116ffb2e3698333ab48e1435f99a6f037b39308f84e62e4fd24888017f06cMTP / speculative {"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --fit-margin 7168env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": false, "context": true, "speed": false}lab_decode_c1_tps 10.2 tok/s lab_prefill_tps 166 tok/s "context-gate prompt tokens / total request seconds"mtp_acceptance_rate 0.5137 fraction {"draft_tokens": 3366, "accepted_tokens": 1729, "source": "server log delta"}mtp_accepted_tps 10.2 tok/s "output tok/s including accepted draft tokens"
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 151350a4d65da543200768e08741f6eb5409c3ed dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T091243Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012-lab.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012
20261004T092615Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012-ttft ✓ pass TTFT T22c eligibility: none
Config ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012 Started 2026-10-04 19:26 AEST Finished 2026-10-04 19:38 AEST (11.7 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab ! TTFT stability ✕
Launch recorded by tools/run.py (artifacts/runs/20261004T091154Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --fit-margin 7168 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168.json sha256 cf6116ffb2e3698333ab48e1435f99a6f037b39308f84e62e4fd24888017f06cMTP / speculative {"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --fit-margin 7168env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 248 tok/s {"prompt_tokens": [10798, 10674, 10665], "ttft_s": [43.785, 43.033, 42.927]}ttft_prefill_32k_tps 224 tok/s {"prompt_tokens": [42970, 42867, 42822], "ttft_s": [191.654, 191.974, 191.365]}ttft_prefill_tps 248 tok/s {"prompt_tokens": [10798, 10674, 10665], "ttft_s": [43.785, 43.033, 42.927]}
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 42d955c4fa51a99398f09218c4a4c133bcf18938 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T092615Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012-ttft.json
Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012
20261004T093800Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012-stability ✕ crash stability T22c eligibility: none failed stage: measure
Config ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012 Started 2026-10-04 19:38 AEST Finished 2026-10-04 19:48 AEST (10.1 min) Exit status 0 Notes stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: engine connection failures (crash/OOM); 9 tool-call parse failure(s); decode floor not verifiable (too little generation)
Other runs of this config: load lab ! TTFT stability ✕
Launch recorded by tools/run.py (artifacts/runs/20261004T091154Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012/docker-run.txt) Copy
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --fit-margin 7168 Image local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Digest sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1Engine ik_llama.cpp Engine source https://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678 Launch file launches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168.json sha256 cf6116ffb2e3698333ab48e1435f99a6f037b39308f84e62e4fd24888017f06cMTP / speculative {"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}argv /app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --fit-margin 7168env CUDA_DEVICE_ORDER=PCI_BUS_IDCUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF @66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json
Metrics context.configured 131,072 tokens context.occupied_max 1,641 tokens context.headroom_min 129,336 tokens duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 0 calls {"per_min": 0.0, "requests": 9}parse_failures 9 responses {"wire": {"unparsed_markup": 9}, "harness_format_errors": {"no_tool_call": 9}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 9 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}decode_window_min_tps unavailable only 41.172 s of generation; need > 120 s for one window decode_whole_run_tps 15.6 tok/s {"generation_seconds": 41.172, "generated_tokens": 642.7}tasks_attempted 5 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 0 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom yes {"upstream_connection_failures": 54, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 26.8, "requests": 3, "tool_calls": 0, "completion_tokens": 208, "occupied_max": 1559}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 28.3, "requests": 3, "tool_calls": 0, "completion_tokens": 227, "occupied_max": 1641}, {"attempt": 3, "round": 1, "instance_id": "psf__requests-2317", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 27.1, "requests": 3, "tool_calls": 0, "completion_tokens": 217, "occupied_max": 1609}, {"attempt": 4, "round": 1, "instance_id": "matplotlib__matplotlib-23299", "exit_status": "Error:BadGatewayError: litellm.BadGatewayError: BadGatewayError: OpenAIException - Cannot connect to host 127.0.0.1:18080 ssl:default [Connect call failed ('127.0.0.1', 18080)]", "submitted": false, "resolved": false, "seconds": 343.6}, {"attempt": 5, "round": 1, "instance_id": "scikit-learn__scikit-learn-13439", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 174.2}] listtasks_over_64k [] instance idscompletion_tokens 652 tokens {"finish_reasons": {"stop": 9}}mtp_accepted_tps unavailable not produced (crash/measure) mtp_acceptance_rate unavailable not produced (crash/measure)
Agentic summary: 600 s, 0 tool calls, 9 parse failures, 0 repetition hits, tasks solved 0/5, decode min window – tok/s, whole run 15.6 tok/s.
Cache, layout and host Cache warm Layout current · 3 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 ● NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 ● NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 42d955c4fa51a99398f09218c4a4c133bcf18938 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T093800Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012-stability.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2
20261004T101015Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72 ✓ pass load T13 eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 Started 2026-10-04 20:10 AEST Finished 2026-10-04 20:14 AEST (3.9 min) Exit status None Notes –
Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261004T101015Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451Digest sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6.json sha256 1671d4e04532370bdb34ebf4958454f7b186984f5668d6c90538d416b0bca440MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 236 s host_ram_drop_gb 63.5 GB vram_ready_mib 20,818 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit bdd5d466ac35145adc76b9f18538d7ef1373acce dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T101015Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2
20261004T101411Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72 ✓ pass lab T13 eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 Started 2026-10-04 20:14 AEST Finished 2026-10-04 20:19 AEST (5.4 min) Exit status 0 Notes served=flashnext
Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261004T101015Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451Digest sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6.json sha256 1671d4e04532370bdb34ebf4958454f7b186984f5668d6c90538d416b0bca440MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max 166,667 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}lab_decode_c1_tps 55.6 tok/s lab_prefill_tps 2,433 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit bdd5d466ac35145adc76b9f18538d7ef1373acce dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T101411Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2
20261004T101937Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72 ✓ pass kit sweep T13 eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 Started 2026-10-04 20:19 AEST Finished 2026-10-04 20:34 AEST (15.2 min) Exit status 0 Notes sweep status: DONE
Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261004T101015Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451Digest sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6.json sha256 1671d4e04532370bdb34ebf4958454f7b186984f5668d6c90538d416b0bca440MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run prefill_8k_tps 2,289 tok/s {"min": 2269.6, "max": 2294.4, "n": 3}prefill_16k_tps 2,240 tok/s {"min": 2233.0, "max": 2247.3, "n": 3}prefill_32k_tps 2,678 tok/s {"min": 2672.6, "max": 2681.2, "n": 3}prefill_64k_tps 2,653 tok/s {"min": 2645.1, "max": 2655.6, "n": 3}decode_c1_tps 48.2 tok/s {"min": 47.55, "max": 48.85, "power_w": null}decode_c2_tps 53.1 tok/s {"min": 51.57, "max": 54.71, "power_w": null}decode_c3_tps 55.9 tok/s {"min": 55.26, "max": 56.61, "power_w": null}decode_c4_tps 53.3 tok/s {"min": 51.51, "max": 55.06, "power_w": null}decode_c1_32k_tps 49.9 tok/s {"min": 49.58, "max": 50.2, "power_w": null}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 9c490cfb769bb4bce8cd979c9e96a84128be7bc8 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T101937Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2
20261004T103452Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72 ✓ pass TTFT T13 eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 Started 2026-10-04 20:34 AEST Finished 2026-10-04 20:36 AEST (1.4 min) Exit status 0 Notes headline length 8192; prompts carry random nonces
Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261004T101015Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451Digest sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6.json sha256 1671d4e04532370bdb34ebf4958454f7b186984f5668d6c90538d416b0bca440MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 2,017 tok/s {"prompt_tokens": [13801, 13745, 13788], "ttft_s": [6.842, 6.83, 6.822]}ttft_prefill_32k_tps 2,687 tok/s {"prompt_tokens": [54693, 54557, 54874], "ttft_s": [20.353, 20.312, 20.371]}ttft_prefill_tps 2,017 tok/s {"prompt_tokens": [13801, 13745, 13788], "ttft_s": [6.842, 6.83, 6.822]}
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 9c490cfb769bb4bce8cd979c9e96a84128be7bc8 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T103452Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2
20261004T103614Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72 ✓ pass quality panel T13 eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 Started 2026-10-04 20:36 AEST Finished 2026-10-04 20:36 AEST (0.4 min) Exit status 0 Notes panel=qwen3.8-flash-next-exl3-ref-panel.json rc=0
Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261004T101015Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451Digest sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6.json sha256 1671d4e04532370bdb34ebf4958454f7b186984f5668d6c90538d416b0bca440MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run top1_agreement 0.98867 fraction mean_kl 0.00096647 nats in_band yes "top1>=0.987, KL<=0.0013"
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 9c490cfb769bb4bce8cd979c9e96a84128be7bc8 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T103614Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2
20261004T103641Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72 ! fail stability T13 eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 Started 2026-10-04 20:36 AEST Finished 2026-10-04 20:47 AEST (10.5 min) Exit status 0 Notes stability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 67 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 37.939 < floor 50.0
Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261004T101015Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451Digest sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6.json sha256 1671d4e04532370bdb34ebf4958454f7b186984f5668d6c90538d416b0bca440MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max 36,327 tokens context.headroom_min 168,259 tokens duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 67 calls {"per_min": 6.7, "requests": 46}parse_failures 0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 0 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /v1/tokenize", "saved": "artifacts: proxy/repetition/"}decode_window_min_tps 37.9 tok/s window 1: 39.97 tok/s window 2: 43.13 tok/s window 3: 42.36 tok/s window 4: 46.42 tok/s window 5: 43.39 tok/s window 6: 42.38 tok/s min 39.97 · max 46.42 tok/s {"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 313, "min_at_generation_s": 87.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}decode_whole_run_tps 44.1 tok/s {"generation_seconds": 432.859, "generated_tokens": 19093.9}tasks_attempted 2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 1 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom no {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 216.5, "requests": 17, "tool_calls": 29, "completion_tokens": 6657, "occupied_max": 13921}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 383.5, "requests": 30, "tool_calls": 38, "completion_tokens": 12541, "occupied_max": 36327}] listtasks_over_64k [] instance idscompletion_tokens 19,198 tokens {"finish_reasons": {"tool_calls": 46}}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Agentic summary: 600 s, 67 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 37.9 tok/s, whole run 44.1 tok/s.
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 9c490cfb769bb4bce8cd979c9e96a84128be7bc8 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T103641Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2
20261004T104713Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72 ✓ pass lab T13 eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 Started 2026-10-04 20:47 AEST Finished 2026-10-04 20:52 AEST (5.3 min) Exit status 0 Notes served=flashnext
Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261004T101015Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451Digest sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6.json sha256 1671d4e04532370bdb34ebf4958454f7b186984f5668d6c90538d416b0bca440MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max 166,667 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}lab_decode_c1_tps 54.5 tok/s lab_prefill_tps 2,596 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 9c490cfb769bb4bce8cd979c9e96a84128be7bc8 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T104713Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-2.05bpw-offload-200k-rtx-3090-24gb-gpu2
20261004T105250Z-sglang-qwen3.8-flash-next-exl3-2.05bpw-offload-200k-rtx-3090-24gb-gpu2-load ✕ crash load T15 eligibility: none failed stage: load
Config sglang-qwen3.8-flash-next-exl3-2.05bpw-offload-200k-rtx-3090-24gb-gpu2 Started 2026-10-04 20:52 AEST Finished 2026-10-04 20:55 AEST (3.0 min) Exit status 0 Notes _glue/offload_moe_method.py", line 86, in get_runtime
store = HostExpertStore(model_dir, layers, num_experts, hidden, inter, bits, prefix=prefix)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/trellis-serve/cuda/src/sglang_exl3/kernels/offload_store.py", line 62, in __init__
raise NotImplementedError("HostExpertStore: fast relayout implemented for K=3 experts only")
NotImplementedError: HostExpertStore: fast relayout implemented for K=3 experts only
[2026-10-04 10:55:46] Received sigquit from a child process. It usually means the child failed.
[2026-10-04 10:55:46] kill_process_tree called: parent_pid=1, include_parent=True, pid=1
[2026-10-04 10:55:46] kill_process_tree called: parent_pid=1, include_parent=False, pid=1
Launch recorded by tools/run.py (artifacts/runs/20261004T105250Z-lab-sglang-qwen3.8-flash-next-exl3-2.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-2.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-2.05bpw_h4_ng4 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-2.05bpw_h4_ng4:/models/turboderp-Qwen3.8-Flash-Next-exl3-2.05bpw_h4_ng4:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-2.05bpw_h4_ng4 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3Digest sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/1e72b250cb529c6d81bf11b88740d9a0b22f62de Launch file launches/sglang-qwen3.8-flash-next-exl3-2.05bpw-offload-200k-rtx-3090-24gb.json sha256 0d170921f6e46f05cad846a1404d158b6e85e4fffe18eeec9a2d7d70ec79a969MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-2.05bpw_h4_ng4 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-2.05bpw_h4_ng4SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @65c895314393431c09050b2e04e250836b3a6eb4 · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-2.05bpw_h4_ng4.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds unavailable not produced (crash/load) host_ram_drop_gb unavailable not produced (crash/load) vram_ready_mib unavailable not produced (crash/load)
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit df0d10f31ae46a27f371fda8e1f15e40dfa878f0 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T105250Z-sglang-qwen3.8-flash-next-exl3-2.05bpw-offload-200k-rtx-3090-24gb-gpu2-load.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-4.05bpw-offload-200k-rtx-3090-24gb-gpu2
20261004T105602Z-sglang-qwen3.8-flash-next-exl3-4.05bpw-offload-200k-rtx-3090-24gb-gpu2-load ✕ crash load T15 eligibility: none failed stage: load
Config sglang-qwen3.8-flash-next-exl3-4.05bpw-offload-200k-rtx-3090-24gb-gpu2 Started 2026-10-04 20:56 AEST Finished 2026-10-04 20:56 AEST (0.9 min) Exit status 0 Notes ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/trellis-serve/cuda/src/sglang_exl3/offload/ngram_nvme.py", line 116, in __init__
t = parse_table(model_path)
^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/trellis-serve/cuda/src/sglang_exl3/offload/ngram_nvme.py", line 94, in parse_table
raise ValueError(f"{path}: shard_N.trellis tensors missing or not contiguous")
ValueError: /models/turboderp-Qwen3.8-Flash-Next-exl3-4.05bpw_h6_ng6/ngram_embedding.safetensors: shard_N.trellis tensors missing or not contiguous
[2026-10-04 10:56:54] Received sigquit from a child process. It usually means the child failed.
[2026-10-04 10:56:54] kill_process_tree called: parent_pid=1, include_parent=True, pid=1
[2026-10-04 10:56:54] kill_process_tree called: parent_pid=1, include_parent=False, pid=1
Launch recorded by tools/run.py (artifacts/runs/20261004T105602Z-lab-sglang-qwen3.8-flash-next-exl3-4.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-4.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-4.05bpw_h6_ng6 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-4.05bpw_h6_ng6:/models/turboderp-Qwen3.8-Flash-Next-exl3-4.05bpw_h6_ng6:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-4.05bpw_h6_ng6 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3Digest sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/1e72b250cb529c6d81bf11b88740d9a0b22f62de Launch file launches/sglang-qwen3.8-flash-next-exl3-4.05bpw-offload-200k-rtx-3090-24gb.json sha256 ce41d62d49f95e441684dbb7afcbfd2cd54c0cdf49476dda70cda1eca5bc7928MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-4.05bpw_h6_ng6 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-4.05bpw_h6_ng6SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @55a732e0c4c3d4614bc42b68493bb930d9b02c0a · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-4.05bpw_h6_ng6.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds unavailable not produced (crash/load) host_ram_drop_gb unavailable not produced (crash/load) vram_ready_mib unavailable not produced (crash/load)
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit bd62e89c96f97db97945fd89be01017a19fea7c3 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T105602Z-sglang-qwen3.8-flash-next-exl3-4.05bpw-offload-200k-rtx-3090-24gb-gpu2-load.json
Runs / GLM-5.3-Flash / exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2
20261004T105700Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-load ✓ pass load T23 eligibility: none
Config exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 Started 2026-10-04 20:57 AEST Finished 2026-10-04 21:02 AEST (5.0 min) Exit status None Notes –
Other runs of this config: load load lab ! lab ! TTFT ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T105700Z-lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30000 --shm-size 16g -e GLM53_MODE=exact -e GLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -e GLM53_MODEL_DOWNLOAD=0 -e HF_HUB_OFFLINE=1 -v ~/models/local-ai-exp/turboderp-GLM-5.3-Flash-exl3-3.05bpw:/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw:ro --entrypoint /opt/glm53/docker/entrypoint.sh ghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flash Image ghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551dDigest sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551dEngine exllamav3 Engine source https://github.com/0xSero/glm53-flash-offload/tree/fe97bcf846347d035859a0610cebbb707fa36caa Launch file launches/exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x.json sha256 b679beb9863b299d0b6e44b1269b954e5292fb8039caa75a0e4c0d34e6cb17faMTP / speculative off argv /opt/glm53/docker/entrypoint.sh -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flashenv GLM53_MODE=exactGLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpwGLM53_MODEL_DOWNLOAD=0HF_HUB_OFFLINE=1
turboderp/GLM-5.3-Flash-exl3 @332ab457b709b7ba30dd9a448be5de03b80a7ac9 · manifest models/turboderp-GLM-5.3-Flash-exl3-3.05bpw.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 300 s host_ram_drop_gb 120 GB vram_ready_mib 23,322 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit d36a5ce7e63d3abd393f2511b2a55930e67bf549 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T105700Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-load.json
Runs / GLM-5.3-Flash / exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2
20261004T110201Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-lab ! fail lab T23 eligibility: none failed stage: gates
Config exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 Started 2026-10-04 21:02 AEST Finished 2026-10-04 21:12 AEST (11.0 min) Exit status 0 Notes served=glm-5.3-flash
Other runs of this config: load load lab ! lab ! TTFT ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T105700Z-lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30000 --shm-size 16g -e GLM53_MODE=exact -e GLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -e GLM53_MODEL_DOWNLOAD=0 -e HF_HUB_OFFLINE=1 -v ~/models/local-ai-exp/turboderp-GLM-5.3-Flash-exl3-3.05bpw:/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw:ro --entrypoint /opt/glm53/docker/entrypoint.sh ghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flash Image ghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551dDigest sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551dEngine exllamav3 Engine source https://github.com/0xSero/glm53-flash-offload/tree/fe97bcf846347d035859a0610cebbb707fa36caa Launch file launches/exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x.json sha256 b679beb9863b299d0b6e44b1269b954e5292fb8039caa75a0e4c0d34e6cb17faMTP / speculative off argv /opt/glm53/docker/entrypoint.sh -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flashenv GLM53_MODE=exactGLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpwGLM53_MODEL_DOWNLOAD=0HF_HUB_OFFLINE=1
turboderp/GLM-5.3-Flash-exl3 @332ab457b709b7ba30dd9a448be5de03b80a7ac9 · manifest models/turboderp-GLM-5.3-Flash-exl3-3.05bpw.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 7.40 tok/s lab_prefill_tps 739 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit d36a5ce7e63d3abd393f2511b2a55930e67bf549 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T110201Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-lab.json
Runs / GLM-5.3-Flash / exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2
20261004T111300Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-ttft ! fail TTFT T23 eligibility: none failed stage: invalid measurement
Config exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 Started 2026-10-04 21:13 AEST Finished 2026-10-04 21:17 AEST (4.2 min) Exit status 0 Notes INVALIDATED: TTFT measured to an empty role chunk the exllamav3 server sends before prefill (ttft.py v1 bug); re-measured with v2. headline length 8192; prompts carry random nonces
Other runs of this config: load load lab ! lab ! TTFT ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T105700Z-lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30000 --shm-size 16g -e GLM53_MODE=exact -e GLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -e GLM53_MODEL_DOWNLOAD=0 -e HF_HUB_OFFLINE=1 -v ~/models/local-ai-exp/turboderp-GLM-5.3-Flash-exl3-3.05bpw:/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw:ro --entrypoint /opt/glm53/docker/entrypoint.sh ghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flash Image ghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551dDigest sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551dEngine exllamav3 Engine source https://github.com/0xSero/glm53-flash-offload/tree/fe97bcf846347d035859a0610cebbb707fa36caa Launch file launches/exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x.json sha256 b679beb9863b299d0b6e44b1269b954e5292fb8039caa75a0e4c0d34e6cb17faMTP / speculative off argv /opt/glm53/docker/entrypoint.sh -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flashenv GLM53_MODE=exactGLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpwGLM53_MODEL_DOWNLOAD=0HF_HUB_OFFLINE=1
turboderp/GLM-5.3-Flash-exl3 @332ab457b709b7ba30dd9a448be5de03b80a7ac9 · manifest models/turboderp-GLM-5.3-Flash-exl3-3.05bpw.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps unavailable invalidated: TTFT measured to an empty role chunk the exllamav3 server sends before prefill (ttft.py v1 bug); re-measured with v2 ttft_prefill_32k_tps unavailable invalidated: TTFT measured to an empty role chunk the exllamav3 server sends before prefill (ttft.py v1 bug); re-measured with v2 ttft_prefill_tps unavailable invalidated: TTFT measured to an empty role chunk the exllamav3 server sends before prefill (ttft.py v1 bug); re-measured with v2
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit d36a5ce7e63d3abd393f2511b2a55930e67bf549 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T111300Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-ttft.json
Runs / GLM-5.3-Flash / exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2
20261004T111712Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-stability ! fail stability T23 eligibility: none
Config exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 Started 2026-10-04 21:17 AEST Finished 2026-10-04 21:27 AEST (10.5 min) Exit status 0 Notes stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 21 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 3.356 < floor 15.0
Other runs of this config: load load lab ! lab ! TTFT ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T105700Z-lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30000 --shm-size 16g -e GLM53_MODE=exact -e GLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -e GLM53_MODEL_DOWNLOAD=0 -e HF_HUB_OFFLINE=1 -v ~/models/local-ai-exp/turboderp-GLM-5.3-Flash-exl3-3.05bpw:/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw:ro --entrypoint /opt/glm53/docker/entrypoint.sh ghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flash Image ghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551dDigest sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551dEngine exllamav3 Engine source https://github.com/0xSero/glm53-flash-offload/tree/fe97bcf846347d035859a0610cebbb707fa36caa Launch file launches/exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x.json sha256 b679beb9863b299d0b6e44b1269b954e5292fb8039caa75a0e4c0d34e6cb17faMTP / speculative off argv /opt/glm53/docker/entrypoint.sh -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flashenv GLM53_MODE=exactGLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpwGLM53_MODEL_DOWNLOAD=0HF_HUB_OFFLINE=1
turboderp/GLM-5.3-Flash-exl3 @332ab457b709b7ba30dd9a448be5de03b80a7ac9 · manifest models/turboderp-GLM-5.3-Flash-exl3-3.05bpw.json
Metrics context.configured 131,072 tokens context.occupied_max 4,589 tokens context.headroom_min 126,405 tokens duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 21 calls {"per_min": 2.1, "requests": 14}parse_failures 0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 0 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}decode_window_min_tps 3.36 tok/s window 1: 6.08 tok/s window 2: 5.96 tok/s min 5.96 · max 6.08 tok/s {"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 99, "min_at_generation_s": 101.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}decode_whole_run_tps 6.56 tok/s {"generation_seconds": 218.025, "generated_tokens": 1431.1}tasks_attempted 2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 1 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom no {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 385.2, "requests": 11, "tool_calls": 16, "completion_tokens": 1605, "occupied_max": 4589}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 214.8, "requests": 4, "tool_calls": 5, "completion_tokens": 179, "occupied_max": 2222}] listtasks_over_64k [] instance idscompletion_tokens 1,784 tokens {"finish_reasons": {"tool_calls": 14}}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Agentic summary: 600 s, 21 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 3.36 tok/s, whole run 6.56 tok/s.
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit d36a5ce7e63d3abd393f2511b2a55930e67bf549 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T111712Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-stability.json
Runs / GLM-5.3-Flash / exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2
20261004T112744Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-lab ! fail lab T23 eligibility: none failed stage: gates
Config exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 Started 2026-10-04 21:27 AEST Finished 2026-10-04 22:05 AEST (37.9 min) Exit status 0 Notes served=glm-5.3-flash
Other runs of this config: load load lab ! lab ! TTFT ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T105700Z-lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30000 --shm-size 16g -e GLM53_MODE=exact -e GLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -e GLM53_MODEL_DOWNLOAD=0 -e HF_HUB_OFFLINE=1 -v ~/models/local-ai-exp/turboderp-GLM-5.3-Flash-exl3-3.05bpw:/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw:ro --entrypoint /opt/glm53/docker/entrypoint.sh ghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flash Image ghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551dDigest sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551dEngine exllamav3 Engine source https://github.com/0xSero/glm53-flash-offload/tree/fe97bcf846347d035859a0610cebbb707fa36caa Launch file launches/exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x.json sha256 b679beb9863b299d0b6e44b1269b954e5292fb8039caa75a0e4c0d34e6cb17faMTP / speculative off argv /opt/glm53/docker/entrypoint.sh -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flashenv GLM53_MODE=exactGLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpwGLM53_MODEL_DOWNLOAD=0HF_HUB_OFFLINE=1
turboderp/GLM-5.3-Flash-exl3 @332ab457b709b7ba30dd9a448be5de03b80a7ac9 · manifest models/turboderp-GLM-5.3-Flash-exl3-3.05bpw.json
Metrics context.configured 131,072 tokens context.occupied_max 93,437 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}lab_decode_c1_tps 7.50 tok/s lab_prefill_tps 720 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit d36a5ce7e63d3abd393f2511b2a55930e67bf549 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T112744Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-lab.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
20261004T120556Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ✓ pass load T13 eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 Started 2026-10-04 22:05 AEST Finished 2026-10-04 22:10 AEST (4.3 min) Exit status None Notes –
Other runs of this config: load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261004T120556Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6.json sha256 cbcfae235adc9e4c7988a047150225a710b8e9ed2b2370701c5d2a9820c28a01MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 260 s host_ram_drop_gb 65.9 GB vram_ready_mib 20,822 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit dcbaa681afe8032d6e59244d27f56918b1e7bd20 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T120556Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
20261004T121018Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ✓ pass lab T13 eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 Started 2026-10-04 22:10 AEST Finished 2026-10-04 22:15 AEST (4.7 min) Exit status 0 Notes served=flashnext
Other runs of this config: load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261004T120556Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6.json sha256 cbcfae235adc9e4c7988a047150225a710b8e9ed2b2370701c5d2a9820c28a01MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max 166,667 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}lab_decode_c1_tps 55.8 tok/s lab_prefill_tps 2,419 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit dcbaa681afe8032d6e59244d27f56918b1e7bd20 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T121018Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
20261004T121501Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ✓ pass kit sweep T13 eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 Started 2026-10-04 22:15 AEST Finished 2026-10-04 22:30 AEST (15.7 min) Exit status 0 Notes sweep status: DONE
Other runs of this config: load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261004T120556Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6.json sha256 cbcfae235adc9e4c7988a047150225a710b8e9ed2b2370701c5d2a9820c28a01MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run prefill_8k_tps 2,281 tok/s {"min": 2261.8, "max": 2289.8, "n": 3}prefill_16k_tps 2,237 tok/s {"min": 2234.0, "max": 2244.7, "n": 3}prefill_32k_tps 2,678 tok/s {"min": 2670.7, "max": 2679.3, "n": 3}prefill_64k_tps 2,649 tok/s {"min": 2643.2, "max": 2654.6, "n": 3}decode_c1_tps 47.4 tok/s {"min": 46.48, "max": 48.27, "power_w": null}decode_c2_tps 53.1 tok/s {"min": 51.34, "max": 54.86, "power_w": null}decode_c3_tps 60.6 tok/s {"min": 57.47, "max": 63.7, "power_w": null}decode_c4_tps 53.7 tok/s {"min": 52.62, "max": 54.77, "power_w": null}decode_c1_32k_tps 48.5 tok/s {"min": 48.52, "max": 48.57, "power_w": null}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit dcbaa681afe8032d6e59244d27f56918b1e7bd20 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T121501Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
20261004T123041Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ✓ pass TTFT T13 eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 Started 2026-10-04 22:30 AEST Finished 2026-10-04 22:32 AEST (1.4 min) Exit status 0 Notes headline length 8192; prompts carry random nonces; TTFT = first chunk with generated output (v2)
Other runs of this config: load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261004T120556Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6.json sha256 cbcfae235adc9e4c7988a047150225a710b8e9ed2b2370701c5d2a9820c28a01MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 2,009 tok/s {"prompt_tokens": [13741, 13755, 13738], "ttft_s": [6.864, 6.84, 6.838]}ttft_prefill_32k_tps 2,681 tok/s {"prompt_tokens": [54658, 54711, 54684], "ttft_s": [20.392, 20.395, 20.396]}ttft_prefill_tps 2,009 tok/s {"prompt_tokens": [13741, 13755, 13738], "ttft_s": [6.864, 6.84, 6.838]}
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit dcbaa681afe8032d6e59244d27f56918b1e7bd20 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T123041Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
20261004T123204Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ✓ pass quality panel T13 eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 Started 2026-10-04 22:32 AEST Finished 2026-10-04 22:32 AEST (0.4 min) Exit status 0 Notes panel=qwen3.8-flash-next-exl3-ref-panel.json rc=0
Other runs of this config: load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261004T120556Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6.json sha256 cbcfae235adc9e4c7988a047150225a710b8e9ed2b2370701c5d2a9820c28a01MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run top1_agreement 0.98967 fraction mean_kl 0.00097327 nats in_band yes "top1>=0.987, KL<=0.0013"
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit dcbaa681afe8032d6e59244d27f56918b1e7bd20 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T123204Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
20261004T123230Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ✕ crash stability T13 eligibility: none failed stage: measure
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 Started 2026-10-04 22:32 AEST Finished 2026-10-04 22:42 AEST (10.1 min) Exit status 0 Notes stability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: engine connection failures (crash/OOM); decode floor not verifiable (too little generation)
Other runs of this config: load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261004T120556Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6.json sha256 cbcfae235adc9e4c7988a047150225a710b8e9ed2b2370701c5d2a9820c28a01MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no successful request context.headroom_min unavailable configured context unknown (set CONFIGURED_CTX) duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 0 calls {"per_min": 0.0, "requests": 0}parse_failures 0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 0 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": null, "saved": "artifacts: proxy/repetition/"}decode_window_min_tps unavailable only 0.0 s of generation; need > 120 s for one window decode_whole_run_tps unavailable no generation recorded tasks_attempted 3 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 0 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom yes {"upstream_connection_failures": 77, "upstream_http_errors": 1, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Error:BadGatewayError: litellm.BadGatewayError: BadGatewayError: OpenAIException - Cannot connect to host 127.0.0.1:18080 ssl:default [Connect call failed ('127.0.0.1', 18080)]", "submitted": false, "resolved": false, "seconds": 267.3}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "Error:BadGatewayError: litellm.BadGatewayError: BadGatewayError: OpenAIException - Cannot connect to host 127.0.0.1:18080 ssl:default [Connect call failed ('127.0.0.1', 18080)]", "submitted": false, "resolved": false, "seconds": 261.8}, {"attempt": 3, "round": 1, "instance_id": "psf__requests-2317", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 71.0}] listtasks_over_64k [] instance idscompletion_tokens 0 tokens {"finish_reasons": {}}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Agentic summary: 600 s, 0 tool calls, 0 parse failures, 0 repetition hits, tasks solved 0/3, decode min window – tok/s, whole run – tok/s.
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit dcbaa681afe8032d6e59244d27f56918b1e7bd20 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T123230Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2
20261004T124252Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ✓ pass load T14a eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2 Started 2026-10-04 22:42 AEST Finished 2026-10-04 22:47 AEST (4.3 min) Exit status None Notes –
Other runs of this config: load lab ✕ TTFT ✕
Launch recorded by tools/run.py (artifacts/runs/20261004T124252Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=6.0 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6.json sha256 90d5526c99bf2c3b87f25e51276eed60130a38e3a34bc2991695048a8bb33eb9MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=6.0SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 257 s host_ram_drop_gb 65.3 GB vram_ready_mib 21,706 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit db208909a5ae27d314ff09d53d7cf803d67fda9c dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T124252Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2
20261004T124710Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ✕ crash lab T14a eligibility: none failed stage: harness
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2 Started 2026-10-04 22:47 AEST Finished 2026-10-04 22:48 AEST (1.7 min) Exit status 1 Notes –
Other runs of this config: load lab ✕ TTFT ✕
Launch recorded by tools/run.py (artifacts/runs/20261004T124252Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=6.0 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6.json sha256 90d5526c99bf2c3b87f25e51276eed60130a38e3a34bc2991695048a8bb33eb9MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=6.0SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run gates_passed unavailable not produced (crash/harness) lab_decode_c1_tps unavailable not produced (crash/harness) lab_prefill_tps unavailable not produced (crash/harness) mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit db208909a5ae27d314ff09d53d7cf803d67fda9c dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T124710Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2
20261004T124852Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ✕ crash TTFT T14a eligibility: none failed stage: measure
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2 Started 2026-10-04 22:48 AEST Finished 2026-10-04 22:49 AEST (0.3 min) Exit status 1 Notes –
Other runs of this config: load lab ✕ TTFT ✕
Launch recorded by tools/run.py (artifacts/runs/20261004T124252Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=6.0 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6.json sha256 90d5526c99bf2c3b87f25e51276eed60130a38e3a34bc2991695048a8bb33eb9MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=6.0SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_tps unavailable not produced (crash/measure)
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit db208909a5ae27d314ff09d53d7cf803d67fda9c dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T124852Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / GLM-5.3-Flash / exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2
20261004T124925Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-load ✓ pass load T23 eligibility: none
Config exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 Started 2026-10-04 22:49 AEST Finished 2026-10-04 22:54 AEST (4.6 min) Exit status None Notes –
Other runs of this config: load load lab ! lab ! TTFT ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T124925Z-lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30000 --shm-size 16g -e GLM53_MODE=exact -e GLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -e GLM53_MODEL_DOWNLOAD=0 -e HF_HUB_OFFLINE=1 -v ~/models/local-ai-exp/turboderp-GLM-5.3-Flash-exl3-3.05bpw:/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw:ro --entrypoint /opt/glm53/docker/entrypoint.sh ghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flash Image ghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551dDigest sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551dEngine exllamav3 Engine source https://github.com/0xSero/glm53-flash-offload/tree/fe97bcf846347d035859a0610cebbb707fa36caa Launch file launches/exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x.json sha256 b679beb9863b299d0b6e44b1269b954e5292fb8039caa75a0e4c0d34e6cb17faMTP / speculative off argv /opt/glm53/docker/entrypoint.sh -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flashenv GLM53_MODE=exactGLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpwGLM53_MODEL_DOWNLOAD=0HF_HUB_OFFLINE=1
turboderp/GLM-5.3-Flash-exl3 @332ab457b709b7ba30dd9a448be5de03b80a7ac9 · manifest models/turboderp-GLM-5.3-Flash-exl3-3.05bpw.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 276 s host_ram_drop_gb 121 GB vram_ready_mib 23,322 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit e2424ec92249689799954b79c057d915d7f402f1 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T124925Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-load.json
Runs / GLM-5.3-Flash / exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2
20261004T125401Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-ttft ✓ pass TTFT T23 eligibility: none
Config exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 Started 2026-10-04 22:54 AEST Finished 2026-10-04 22:58 AEST (4.2 min) Exit status 0 Notes headline length 8192; prompts carry random nonces; TTFT = first chunk with generated output (v2)
Other runs of this config: load load lab ! lab ! TTFT ! TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T124925Z-lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2/docker-run.txt) Copy
docker run -d --name lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30000 --shm-size 16g -e GLM53_MODE=exact -e GLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -e GLM53_MODEL_DOWNLOAD=0 -e HF_HUB_OFFLINE=1 -v ~/models/local-ai-exp/turboderp-GLM-5.3-Flash-exl3-3.05bpw:/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw:ro --entrypoint /opt/glm53/docker/entrypoint.sh ghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flash Image ghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551dDigest sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551dEngine exllamav3 Engine source https://github.com/0xSero/glm53-flash-offload/tree/fe97bcf846347d035859a0610cebbb707fa36caa Launch file launches/exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x.json sha256 b679beb9863b299d0b6e44b1269b954e5292fb8039caa75a0e4c0d34e6cb17faMTP / speculative off argv /opt/glm53/docker/entrypoint.sh -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flashenv GLM53_MODE=exactGLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpwGLM53_MODEL_DOWNLOAD=0HF_HUB_OFFLINE=1
turboderp/GLM-5.3-Flash-exl3 @332ab457b709b7ba30dd9a448be5de03b80a7ac9 · manifest models/turboderp-GLM-5.3-Flash-exl3-3.05bpw.json
Metrics context.configured 131,072 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 483 tok/s {"prompt_tokens": [10753, 10793, 10717], "ttft_s": [19.173, 22.332, 25.068]}ttft_prefill_32k_tps 688 tok/s {"prompt_tokens": [42865, 42692, 42811], "ttft_s": [62.271, 62.501, 60.752]}ttft_prefill_tps 483 tok/s {"prompt_tokens": [10753, 10793, 10717], "ttft_s": [19.173, 22.332, 25.068]}
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit e2424ec92249689799954b79c057d915d7f402f1 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/glm-5.3-flash/20261004T125401Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-ttft.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
20261004T125827Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ✓ pass load T14a eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 Started 2026-10-04 22:58 AEST Finished 2026-10-04 23:02 AEST (4.4 min) Exit status None Notes –
Other runs of this config: load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261004T125827Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6.json sha256 cbcfae235adc9e4c7988a047150225a710b8e9ed2b2370701c5d2a9820c28a01MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 264 s host_ram_drop_gb 65.7 GB vram_ready_mib 20,844 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 9229192eace16fc1782f479846e0194aea593ba9 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T125827Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
20261004T130251Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ! fail stability T14a eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 Started 2026-10-04 23:02 AEST Finished 2026-10-04 23:13 AEST (10.5 min) Exit status 0 Notes stability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 71 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 39.676 < floor 50.0
Other runs of this config: load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261004T125827Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6.json sha256 cbcfae235adc9e4c7988a047150225a710b8e9ed2b2370701c5d2a9820c28a01MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max 25,309 tokens context.headroom_min 178,942 tokens duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 71 calls {"per_min": 7.1, "requests": 47}parse_failures 0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 0 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /v1/tokenize", "saved": "artifacts: proxy/repetition/"}decode_window_min_tps 39.7 tok/s window 1: 44.31 tok/s window 2: 43.52 tok/s window 3: 43.73 tok/s window 4: 44.87 tok/s window 5: 41.92 tok/s min 41.92 · max 44.87 tok/s {"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 254, "min_at_generation_s": 267.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}decode_whole_run_tps 43.3 tok/s {"generation_seconds": 373.189, "generated_tokens": 16176.9}tasks_attempted 2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 1 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom no {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 245.6, "requests": 21, "tool_calls": 31, "completion_tokens": 5704, "occupied_max": 13986}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 354.4, "requests": 27, "tool_calls": 40, "completion_tokens": 10614, "occupied_max": 25309}] listtasks_over_64k [] instance idscompletion_tokens 16,318 tokens {"finish_reasons": {"tool_calls": 47}}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Agentic summary: 600 s, 71 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 39.7 tok/s, whole run 43.3 tok/s.
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 9229192eace16fc1782f479846e0194aea593ba9 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T130251Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
20261004T131323Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ✓ pass lab T14a eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 Started 2026-10-04 23:13 AEST Finished 2026-10-04 23:22 AEST (8.9 min) Exit status 0 Notes served=flashnext
Other runs of this config: load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261004T125827Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6.json sha256 cbcfae235adc9e4c7988a047150225a710b8e9ed2b2370701c5d2a9820c28a01MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max 166,667 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}lab_decode_c1_tps 57.4 tok/s lab_prefill_tps 2,537 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 9229192eace16fc1782f479846e0194aea593ba9 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T131323Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2
20261004T132232Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ✓ pass load T14a eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2 Started 2026-10-04 23:22 AEST Finished 2026-10-04 23:26 AEST (4.2 min) Exit status None Notes –
Other runs of this config: load lab lab TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T132232Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 1 --cuda-graph-max-bs-decode 1 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1.json sha256 e76072403dd6ef64022ca8b6ccb17b7e7d25f20aa7c17b8287c7534158eee184MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 1 --cuda-graph-max-bs-decode 1 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 254 s host_ram_drop_gb 64.4 GB vram_ready_mib 20,744 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 97294db30283c264f2c58088bb36bfb1619ba3de dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T132232Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2
20261004T132647Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ✓ pass lab T14a eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2 Started 2026-10-04 23:26 AEST Finished 2026-10-04 23:31 AEST (5.0 min) Exit status 0 Notes served=flashnext
Other runs of this config: load lab lab TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T132232Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 1 --cuda-graph-max-bs-decode 1 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1.json sha256 e76072403dd6ef64022ca8b6ccb17b7e7d25f20aa7c17b8287c7534158eee184MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 1 --cuda-graph-max-bs-decode 1 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max 166,667 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}lab_decode_c1_tps 55.6 tok/s lab_prefill_tps 2,430 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 97294db30283c264f2c58088bb36bfb1619ba3de dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T132647Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2
20261004T133146Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ✓ pass TTFT T14a eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2 Started 2026-10-04 23:31 AEST Finished 2026-10-04 23:33 AEST (1.4 min) Exit status 0 Notes headline length 8192; prompts carry random nonces; TTFT = first chunk with generated output (v2)
Other runs of this config: load lab lab TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T132232Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 1 --cuda-graph-max-bs-decode 1 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1.json sha256 e76072403dd6ef64022ca8b6ccb17b7e7d25f20aa7c17b8287c7534158eee184MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 1 --cuda-graph-max-bs-decode 1 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 2,014 tok/s {"prompt_tokens": [13689, 13759, 13721], "ttft_s": [6.848, 6.829, 6.812]}ttft_prefill_32k_tps 2,692 tok/s {"prompt_tokens": [54692, 54821, 54730], "ttft_s": [20.316, 20.367, 20.356]}ttft_prefill_tps 2,014 tok/s {"prompt_tokens": [13689, 13759, 13721], "ttft_s": [6.848, 6.829, 6.812]}
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 97294db30283c264f2c58088bb36bfb1619ba3de dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T133146Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2
20261004T133309Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ! fail stability T14a eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2 Started 2026-10-04 23:33 AEST Finished 2026-10-04 23:43 AEST (10.5 min) Exit status 0 Notes stability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 57 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 40.08 < floor 50.0
Other runs of this config: load lab lab TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T132232Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 1 --cuda-graph-max-bs-decode 1 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1.json sha256 e76072403dd6ef64022ca8b6ccb17b7e7d25f20aa7c17b8287c7534158eee184MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 1 --cuda-graph-max-bs-decode 1 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max 29,027 tokens context.headroom_min 174,880 tokens duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 57 calls {"per_min": 5.7, "requests": 42}parse_failures 0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 0 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /v1/tokenize", "saved": "artifacts: proxy/repetition/"}decode_window_min_tps 40.1 tok/s window 1: 42.01 tok/s window 2: 46.09 tok/s window 3: 40.09 tok/s window 4: 43.71 tok/s min 40.09 · max 46.09 tok/s {"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 193, "min_at_generation_s": 181.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}decode_whole_run_tps 43.1 tok/s {"generation_seconds": 312.161, "generated_tokens": 13447.2}tasks_attempted 2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 1 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom no {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 124.1, "requests": 11, "tool_calls": 18, "completion_tokens": 3257, "occupied_max": 8494}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 475.9, "requests": 32, "tool_calls": 39, "completion_tokens": 10303, "occupied_max": 29027}] listtasks_over_64k [] instance idscompletion_tokens 13,560 tokens {"finish_reasons": {"tool_calls": 42}}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Agentic summary: 600 s, 57 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 40.1 tok/s, whole run 43.1 tok/s.
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 97294db30283c264f2c58088bb36bfb1619ba3de dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T133309Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2
20261004T134340Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ✓ pass lab T14a eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2 Started 2026-10-04 23:43 AEST Finished 2026-10-04 23:48 AEST (4.6 min) Exit status 0 Notes served=flashnext
Other runs of this config: load lab lab TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T132232Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 1 --cuda-graph-max-bs-decode 1 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1.json sha256 e76072403dd6ef64022ca8b6ccb17b7e7d25f20aa7c17b8287c7534158eee184MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 1 --cuda-graph-max-bs-decode 1 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max 166,667 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}lab_decode_c1_tps 54.6 tok/s lab_prefill_tps 2,580 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 97294db30283c264f2c58088bb36bfb1619ba3de dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T134340Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2
20261004T134832Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ✓ pass load T14b eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2 Started 2026-10-04 23:48 AEST Finished 2026-10-04 23:52 AEST (3.6 min) Exit status None Notes –
Other runs of this config: load lab lab TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T134832Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=pinned -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned.json sha256 06f6cc13b0c25dfb2d50a8b515c6240bf1418dfc8b07facf231d3e0847358cb1MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=pinnedSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 215 s host_ram_drop_gb 88.2 GB vram_ready_mib 20,740 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 1b7b92a2ae73d921b9b6ab2ff7cb2059734faf95 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T134832Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2
20261004T135208Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ✓ pass lab T14b eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2 Started 2026-10-04 23:52 AEST Finished 2026-10-04 23:55 AEST (3.7 min) Exit status 0 Notes served=flashnext
Other runs of this config: load lab lab TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T134832Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=pinned -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned.json sha256 06f6cc13b0c25dfb2d50a8b515c6240bf1418dfc8b07facf231d3e0847358cb1MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=pinnedSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max 166,667 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}lab_decode_c1_tps 53.7 tok/s lab_prefill_tps 2,440 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 1b7b92a2ae73d921b9b6ab2ff7cb2059734faf95 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T135208Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2
20261004T135547Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ✓ pass TTFT T14b eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2 Started 2026-10-04 23:55 AEST Finished 2026-10-04 23:57 AEST (1.4 min) Exit status 0 Notes headline length 8192; prompts carry random nonces; TTFT = first chunk with generated output (v2)
Other runs of this config: load lab lab TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T134832Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=pinned -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned.json sha256 06f6cc13b0c25dfb2d50a8b515c6240bf1418dfc8b07facf231d3e0847358cb1MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=pinnedSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run ttft_prefill_8k_tps 2,022 tok/s {"prompt_tokens": [13765, 13739, 13758], "ttft_s": [6.807, 6.819, 6.799]}ttft_prefill_32k_tps 2,700 tok/s {"prompt_tokens": [54781, 54752, 54783], "ttft_s": [20.259, 20.275, 20.301]}ttft_prefill_tps 2,022 tok/s {"prompt_tokens": [13765, 13739, 13758], "ttft_s": [6.807, 6.819, 6.799]}
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 1b7b92a2ae73d921b9b6ab2ff7cb2059734faf95 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T135547Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2
20261004T135709Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ! fail stability T14b eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2 Started 2026-10-04 23:57 AEST Finished 2026-10-05 00:07 AEST (10.5 min) Exit status 0 Notes stability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 72 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 40.915 < floor 50.0
Other runs of this config: load lab lab TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T134832Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=pinned -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned.json sha256 06f6cc13b0c25dfb2d50a8b515c6240bf1418dfc8b07facf231d3e0847358cb1MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=pinnedSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max 39,281 tokens context.headroom_min 165,395 tokens duration_s 600 s {"cap_s": 600.0, "driver_capped": true}tool_calls 72 calls {"per_min": 7.2, "requests": 50}parse_failures 0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}output_truncations 0 responses {"rule": "cut off at max_tokens with no tool call"}agent_format_errors 0 responses {"rule": "agent protocol errors other than malformed tool calls"}repetition_hits 0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /v1/tokenize", "saved": "artifacts: proxy/repetition/"}decode_window_min_tps 40.9 tok/s window 1: 46.31 tok/s window 2: 42.86 tok/s window 3: 45.63 tok/s window 4: 44.46 tok/s window 5: 47.12 tok/s min 42.86 · max 47.12 tok/s {"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 289, "min_at_generation_s": 108.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}decode_whole_run_tps 44.6 tok/s {"generation_seconds": 408.606, "generated_tokens": 18207.3}tasks_attempted 2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}tasks_solved 1 attempts {"eval_errors": 0, "swebench": "5.0.2"}crash_or_oom no {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}gates_after skipped filled by qualify step: the lab run that follows this run (tools/qualify.py) agent_attempts [{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 137.5, "requests": 16, "tool_calls": 22, "completion_tokens": 3550, "occupied_max": 10227}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 462.8, "requests": 34, "tool_calls": 50, "completion_tokens": 14781, "occupied_max": 39281}] listtasks_over_64k [] instance idscompletion_tokens 18,331 tokens {"finish_reasons": {"tool_calls": 50}}mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Agentic summary: 600 s, 72 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 40.9 tok/s, whole run 44.6 tok/s.
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 1b7b92a2ae73d921b9b6ab2ff7cb2059734faf95 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T135709Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2
20261004T140742Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ✓ pass lab T14b eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2 Started 2026-10-05 00:07 AEST Finished 2026-10-05 00:11 AEST (3.7 min) Exit status 0 Notes served=flashnext
Other runs of this config: load lab lab TTFT stability !
Launch recorded by tools/run.py (artifacts/runs/20261004T134832Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=pinned -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned.json sha256 06f6cc13b0c25dfb2d50a8b515c6240bf1418dfc8b07facf231d3e0847358cb1MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=pinnedSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max 166,667 tokens context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}lab_decode_c1_tps 51.3 tok/s lab_prefill_tps 2,596 tok/s "context-gate prompt tokens / total request seconds"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache warm Layout current · 1 card(s) used · display off hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 off 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 1b7b92a2ae73d921b9b6ab2ff7cb2059734faf95 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T140742Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
20261004T141422Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ✓ pass load T13 eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 Started 2026-10-05 00:14 AEST Finished 2026-10-05 00:18 AEST (4.3 min) Exit status None Notes –
Other runs of this config: load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261004T141422Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6.json sha256 cbcfae235adc9e4c7988a047150225a710b8e9ed2b2370701c5d2a9820c28a01MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run load_seconds 257 s host_ram_drop_gb 64.5 GB vram_ready_mib 20,818 MiB
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 742e81596c31f25802bdf11347fc6d9ab1bb8265 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T141422Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
20261004T141839Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe ✓ pass lab T13 eligibility: none
Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 Started 2026-10-05 00:18 AEST Finished 2026-10-05 00:25 AEST (7.0 min) Exit status 0 Notes official lab.py try --on endpoint; run file rtx-3090-24gb.qwen3.8-flash-next.sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb.200k.20261004T142537.json; recipe written only on pass
Other runs of this config: load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel
Launch recorded by tools/run.py (artifacts/runs/20261004T141422Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt) Copy
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder Image ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Digest sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163Engine sglang Engine source https://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a Launch file launches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6.json sha256 cbcfae235adc9e4c7988a047150225a710b8e9ed2b2370701c5d2a9820c28a01MTP / speculative off argv /opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coderenv HF_HUB_OFFLINE=1CUDA_DEVICE_ORDER=PCI_BUS_IDSGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5SGLANG_EXL3_MOE_OFFLOAD=gpu_cacheEXL3_MOE_CPU_THREADS=24SGLANG_EXL3_EXPERT_CACHE_GB=autoSGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26SGLANG_EXL3_KV_BITS=5SGLANG_EXL3_EMBED_HOST=1SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4SGLANG_EXL3_OFFLOAD_FUSED=1SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1SGLANG_EXL3_NGRAM_TIER=nvmeSGLANG_EXL3_VISION=1SGLANG_EXL3_MM_FAST_CPU=1SGLANG_EXL3_VIT_SDPA=1SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3 @69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json
Metrics context.configured 204,800 tokens context.occupied_max unavailable no occupancy measured in this run context.headroom_min unavailable no occupancy measured in this run gates_passed ["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}lab_decode_c1_tps 55.8 tok/s lab_prefill_tps 2,419 tok/s "official lab.py try"mtp_accepted_tps skipped speculative decoding off in this launch mtp_acceptance_rate skipped speculative decoding off in this launch
Cache, layout and host Cache cold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache Layout current · 1 card(s) used · display on hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 cpus_offline 2,18 cpu_boost False cpu_max_mhz 3401 machine_checks_this_boot 6
# Used Name UUID Bus PCIe Display 0 NVIDIA GeForce RTX 3090 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off 1 NVIDIA GeForce RTX 3090 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on 2 ● NVIDIA GeForce RTX 3090 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
Provenance bench commit 742e81596c31f25802bdf11347fc6d9ab1bb8265 dirty tree pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)registry d21258dd744e7c78be28177c90060c6af8e10b7ekit ef883d269e50ecbea290f919f095e2f3ca633b42mini-swe-agent 04d809ceab9df28f9adaed044884180159172930
Artifacts and raw output Raw record: results/qwen3.8-flash-next/20261004T141839Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json
Hardware & methodology Host-to-device bandwidth Run Outcome Layout GPU Bus PCIe H2D D2H Notes 2026-10-03 06:53 AEST ✓ pass current GPU2 GPU-76d3c6af-7f18-69d2-3145-899103de1722 00000000:0C:00.0gen4 x8 13.5 GB/s 13.2 GB/s 1 GiB pinned transfers; full table in detail 2026-10-03 06:53 AEST ✓ pass current GPU0 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d 00000000:04:00.0gen4 x8 6.03 GB/s 6.60 GB/s 1 GiB pinned transfers; full table in detail 2026-10-04 19:11 AEST ✓ pass current GPU1 GPU-c67ac872-3371-a88c-aa92-d969daf6405e 00000000:0B:00.0gen4 x8 13.5 GB/s 13.2 GB/s 1 GiB pinned transfers; full table in detail
NVMe random reads Run Outcome IOPS Throughput Detail Notes 2026-10-03 06:54 AEST ✓ pass 168,600 IOPS 3.35 GB/s "16 KiB, 32 threads"~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5/ngram_embedding.safetensors
Host hostname omarchy-gpu cpu AMD Ryzen 9 5950X 16-Core Processor ram_total_gib 125.7 ram_layout DIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels) kernel 7.2.5-3-omarchy driver 610.57.04 cuda 13.3 cpus_online 0-1,3-17,19-31 0-31 (changed between runs) cpus_offline 2,18 none (changed between runs) machine_checks_this_boot 0 6 (changed between runs) cpu_boost False cpu_max_mhz 3401
Layout GPU UUID Bus PCIe Display current GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8 off current GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8 on current GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8 off
GPUs as seen by the records (PCIe width is the link width at record time).
Dropped configs and why Model Config Failed runs Why (record notes) GLM exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 ! fail lab / gates ! fail TTFT / invalid measurement ! fail stability ! fail lab / gatesserved=glm-5.3-flash INVALIDATED: TTFT measured to an empty role chunk the exllamav3 server sends before prefill (ttft.py v1 bug); re-measured with v2. headline length 8192; prompts carry random nonces stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 21 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 3.356 < floor 15.0 served=glm-5.3-flash GLM glm-llamacpp-iq2m-128k-1x-gpu2 ! fail lab / gates ! fail stability ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 24 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 8.91 < floor 15.0 served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf GLM ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 ! fail lab / gates ! fail stability ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: 47 tool-call parse failure(s); decode window min 9.458 < floor 15.0 served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf GLM ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2 ! fail lab / gates ✕ crash stability / measureserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: engine connection failures (crash/OOM); 9 tool-call parse failure(s); decode floor not verifiable (too little generation) GLM ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2 ! fail lab / gates ✕ crash stability / measureserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: engine connection failures (crash/OOM); 9 tool-call parse failure(s); decode floor not verifiable (too little generation) GLM ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 ! fail lab / gates ! fail stability ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: 68 tool-call parse failure(s); decode window min 13.255 < floor 15.0 served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf GLM ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-gpu012 ✕ crash load / loadche_init: CUDA2 KV buffer size = 570.81 MiB
llama_init_from_model: KV self size = 748.00 MiB, c^KV (q8_0): 748.00 MiB, kv^T: not used
llama_init_from_model: KV self indexer size = 704.00 MiB (f16)
llama_init_from_model: CUDA_Host output buffer size = 0.59 MiB
ggml_backend_cuda_buffer_type_alloc_buffer: allocating 3488.51 MiB on device 0: cudaMalloc failed: out of memory
ggml_gallocr_reserve_n: failed to allocate CUDA0 buffer of size 3657965696
llama_init_from_model: failed to allocate compute buffers
~ggml_backend_cuda_context: have 0 graphs
~ggml_backend_cuda_context: have 0 graphs
~ggml_backend_cuda_context: have 0 graphs
llama_init_from_gpt_params: error: failed to create context with model '/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf'
GLM ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm4096-gpu012 ✕ crash load / load c^KV (q8_0): 68.00 MiB, kv^T: not used
llama_init_from_model: KV self indexer size = 64.00 MiB (f16)
llama_init_from_model: CUDA_Host output buffer size = 0.61 MiB
ggml_backend_cuda_buffer_type_alloc_buffer: allocating 2754.25 MiB on device 0: cudaMalloc failed: out of memory
ggml_gallocr_reserve_n: failed to allocate CUDA0 buffer of size 2888039424
llama_init_from_model: failed to allocate compute buffers
~ggml_backend_cuda_context: have 0 graphs
~ggml_backend_cuda_context: have 0 graphs
~ggml_backend_cuda_context: have 0 graphs
common_speculative_init: failed to create MTP context
srv init: failed to initialize recurrent speculative context
~ggml_backend_cuda_context: have 13 graphs
~ggml_backend_cuda_context: have 14 graphs
~ggml_backend_cuda_context: have 11 graphs
GLM ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012 ! fail lab / gates ✕ crash stability / measureserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: engine connection failures (crash/OOM); 9 tool-call parse failure(s); decode floor not verifiable (too little generation) GLM ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-gpu012 ✕ crash load / loadche_init: CUDA2 KV buffer size = 570.81 MiB
llama_init_from_model: KV self size = 816.00 MiB, c^KV (q8_0): 816.00 MiB, kv^T: not used
llama_init_from_model: KV self indexer size = 768.00 MiB (f16)
llama_init_from_model: CUDA_Host output buffer size = 0.61 MiB
ggml_backend_cuda_buffer_type_alloc_buffer: allocating 3488.51 MiB on device 0: cudaMalloc failed: out of memory
ggml_gallocr_reserve_n: failed to allocate CUDA0 buffer of size 3657965696
llama_init_from_model: failed to allocate compute buffers
~ggml_backend_cuda_context: have 0 graphs
~ggml_backend_cuda_context: have 0 graphs
~ggml_backend_cuda_context: have 0 graphs
llama_init_from_gpt_params: error: failed to create context with model '/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf'
GLM llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 ! fail lab / gates ! fail stability ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 30 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 8.854 < floor 15.0 served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf GLM llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2 ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf GLM llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 ! fail lab / gates ! fail stability ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 27 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: 1 repetition hit(s); decode window min 8.847 < floor 15.0 served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf GLM llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2 ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf GLM llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012 ! fail lab / gates ! fail stability ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 24 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 11.035 < floor 15.0 served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf GLM llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012 ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf GLM llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 ! fail lab / gates ! fail stability ✕ crash lab / harnessserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 26 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 11.388 < floor 15.0 (no notes) GLM llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012 ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf GLM llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 ! fail lab / gates ! fail stability ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 25 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 7.646 < floor 15.0 served=/models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf GLM llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 ! fail lab / gates ! fail stability ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 22 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: 1 repetition hit(s); decode window min 9.855 < floor 15.0 served=/models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf GLM llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012 ! fail lab / gates ! fail stability ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 25 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 7.807 < floor 15.0 served=/models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf GLM llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 ! fail lab / gates ! fail stability ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 19 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 8.8 < floor 15.0 served=/models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf Flash-Next fn-llamacpp-iq4xs-128k-1x-fit-gpu2 ! fail stabilitystability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 49 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: 1 repetition hit(s); decode window min 23.83 < floor 50.0 Flash-Next fn-llamacpp-iq4xs-128k-1x-gpu2 ✕ crash load / loaduser, abort
0.18.994.031 E ggml_backend_cuda_buffer_type_alloc_buffer: allocating 65358.17 MiB on device 0: cudaMalloc failed: out of memory
0.18.994.037 E alloc_tensor_range: failed to allocate CUDA0 buffer of size 68533006336
0.19.139.695 E llama_model_load: error loading model: unable to allocate CUDA0 buffer
0.19.139.702 E llama_model_load_from_file_impl: failed to load model
0.19.139.708 E cmn common_init_: failed to load model '/models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf'
0.19.139.713 E srv load_model: failed to load model, '/models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf'
0.19.139.715 I srv operator(): operator(): cleaning up before exit...
0.19.140.570 E srv llama_server: exiting due to model loading error
Flash-Next fn-trellis-3.05-0xsero-main-x8-gpu2 ! fail stabilitystability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 63 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 38.077 < floor 50.0 Flash-Next llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012 ! fail stabilitystability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 70 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 38.626 < floor 50.0 Flash-Next llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012 ! fail stabilitystability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode floor not verifiable (too little generation); agent driver exited 1 Flash-Next sglang-qwen3.8-flash-next-exl3-2.05bpw-offload-200k-rtx-3090-24gb-gpu2 ✕ crash load / load_glue/offload_moe_method.py", line 86, in get_runtime
store = HostExpertStore(model_dir, layers, num_experts, hidden, inter, bits, prefix=prefix)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/trellis-serve/cuda/src/sglang_exl3/kernels/offload_store.py", line 62, in __init__
raise NotImplementedError("HostExpertStore: fast relayout implemented for K=3 experts only")
NotImplementedError: HostExpertStore: fast relayout implemented for K=3 experts only
[2026-10-04 10:55:46] Received sigquit from a child process. It usually means the child failed.
[2026-10-04 10:55:46] kill_process_tree called: parent_pid=1, include_parent=True, pid=1
[2026-10-04 10:55:46] kill_process_tree called: parent_pid=1, include_parent=False, pid=1
Flash-Next sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2 ! fail stabilitystability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 57 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 40.08 < floor 50.0 Flash-Next sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 ✕ crash stability / measure ! fail stabilitystability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: engine connection failures (crash/OOM); decode floor not verifiable (too little generation) stability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 71 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 39.676 < floor 50.0 Flash-Next sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2 ! fail stabilitystability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 72 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 40.915 < floor 50.0 Flash-Next sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2 ✕ crash lab / harness ✕ crash TTFT / measure(no notes) (no notes) Flash-Next sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 ! fail stabilitystability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 67 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 37.939 < floor 50.0 Flash-Next sglang-qwen3.8-flash-next-exl3-4.05bpw-offload-200k-rtx-3090-24gb-gpu2 ✕ crash load / load ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/trellis-serve/cuda/src/sglang_exl3/offload/ngram_nvme.py", line 116, in __init__
t = parse_table(model_path)
^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/trellis-serve/cuda/src/sglang_exl3/offload/ngram_nvme.py", line 94, in parse_table
raise ValueError(f"{path}: shard_N.trellis tensors missing or not contiguous")
ValueError: /models/turboderp-Qwen3.8-Flash-Next-exl3-4.05bpw_h6_ng6/ngram_embedding.safetensors: shard_N.trellis tensors missing or not contiguous
[2026-10-04 10:56:54] Received sigquit from a child process. It usually means the child failed.
[2026-10-04 10:56:54] kill_process_tree called: parent_pid=1, include_parent=True, pid=1
[2026-10-04 10:56:54] kill_process_tree called: parent_pid=1, include_parent=False, pid=1
All failed and crashed runs (64) Decisions 0001-flashnext-llamacpp-quant.md 0001 — Flash-Next llama.cpp comparison quant (T11)0001 — Flash-Next llama.cpp comparison quant (T11)
Decision: use bartowski/Qwen3.8-Flash-Next-GGUF@928589fdb66c6ff07f22ac561e3fbce76553548f IQ4_XS (97.7 GB) for both the 1-card and the 3-card llama.cpp comparison, on the attested ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312 image. N-gram table (per_layer_token_embd.weight) pinned to CPU with -ot; experts on CPU for the layers that do not fit (--n-cpu-moe N, tuned once per layout), fp16→q8_0 KV only if needed for 128k+.
Why: closest to 4-bit that fits 24 GB VRAM + ~110 GiB RAM on one card (IQ4_XS ≈ 91 GiB); same file serves the 3-card run, so the 1 vs 3 comparison changes only the layout. Q4_K_M (119.6 GB ≈ 111 GiB) would not leave room on 1 card. No MTP (llama.cpp Flash-Next MTP PR #27836 not merged). Not REAP-pruned.
Limits: the kit quality panel speaks the SGLang API only, so this route gets no kit band → it cannot be the Flash-Next headline (SPEC §5); it is a comparison point. Delete after T11 unless it beats trellis.
Outcome after review (GPT-6.1 Sol: AMEND) — accepted
IQ4_XS stays the paired 1 vs 3 card file. Q4_K_M only as a single feasibility screen if IQ4_XS is competitive with trellis (disk is the binding constraint).
The 1 vs 3 pair freezes occupied context, KV format, batching and workload; --n-cpu-moe / tensor split per layout are reported, and any KV change is a separate comparison.
Full digest + shard hashes via tools/models.py; tensor placement of per_layer_token_embd.weight verified from the server log.
Eligibility experiment-only (no kit band). Results kept regardless; weights deleted only after manifests exist.
0001-flashnext-llamacpp-quant.review.md 0001-flashnext-llamacpp-quant.review
AMEND. Keep IQ4_XS as the primary paired comparison, but replace the categorical Q4_K_M rejection with a memory-budget check. The strongest counter-argument is that the decision compares the entire file against host RAM , despite offloading weights to VRAM.
Q4_K_M is not demonstrably excluded. 119.6 GB ≈ 111.4 GiB . If 18–20 GiB of weights reside on the GPU and their host pages are reclaimable, host weight residency becomes roughly 91–93 GiB , leaving 17–19 GiB within the stated 110 GiB RAM budget. Loading peaks, pinned allocations, KV and scratch could consume that margin; the supplied numbers do not prove they will.
Concrete alternative: use IQ4_XS for T11, with one bounded Q4_K_M feasibility screen at the intended context if the allocation budget supports it. Keep any Q4 results as a separate quant comparison. Do not expand into another tuning grid.
Fitting is not evidence of adequate speed. Dual-channel DDR4-3200 has a theoretical 51.2 GB/s ceiling. At 50 tok/s, that permits only 1.024 GB of DRAM traffic per token . If all 6B active LM parameters required fresh ideal 4-bit reads, that is ~3 GB/token and at most 17 tok/s before overhead. Record CPU offload placement and measured decode; this is a conditional bandwidth screen, not a performance prediction.
“Only the layout changes” needs qualification. Independently tuning --n-cpu-moe N compares each layout’s selected configuration, rather than isolating card-count scaling. Freeze occupied context, KV format, batching and workload across the pair; report both CPU-expert allocations and tensor splits. Treat any KV-format switch as a separate comparison.
Complete provenance before launch. Pin the full image digest already supplied in SPEC §6, not just server-cuda12-b11312. Add the exact GGUF shard list and hashes, immutable launches, and the actual tensor-placement evidence for per_layer_token_embd.weight. The proposed -ot setting remains unverified in this review.
The quality limitation is correctly identified, but applies to registry eligibility too. Without the required kit band, this Flash-Next configuration is experiment-only , even if all registry lab gates pass. Neither IQ4_XS nor a speed win establishes quality equivalence.
Replace “delete unless it beats trellis.” Define the comparison at matched occupied context and cache state, and retain results regardless of outcome. Delete rejected weights only after hashes and immutable references are recorded; retain any finalist under §9. A useful three-card result need not beat a single-GPU engine to justify publication.
0002-glm-quant-plan.md 0002 — GLM-5.3-Flash quant plan for llama.cpp / ik_llama.cpp (T20–T22)0002 — GLM-5.3-Flash quant plan for llama.cpp / ik_llama.cpp (T20–T22)
Source: bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 (made with llama.cpp b11279, mainline-compatible; MTP layers kept at Q4_0). Unsloth GGUFs excluded: not loadable on mainline until re-uploaded.
Layout Quants Size Fit reasoning 1 card (24 GB VRAM + ~115 GiB usable RAM) IQ2_M, Q2_K 120.6 / 125.7 GB (112 / 117 GiB) ~21 GiB on GPU, rest in RAM 1 card, stretch IQ3_XXS 138.8 GB (129 GiB) over RAM by ~10 GiB → mmap paging from NVMe; run once, expect slow, record it 3 cards (72 GB VRAM + RAM) IQ3_XXS, IQ4_XS 138.8 / 177.7 GB IQ4_XS ≈ 165 GiB: ~66 GiB VRAM + ~99 GiB RAM; ~4-bit like every accepted GLM recipe
Order (disk is ~750 GB for everything, so sequential): IQ2_M → Q2_K → (IQ3_XXS) → delete the losing 2-bit → IQ4_XS. ik_llama.cpp + MTP runs on the same files.
Quality judged per SPEC §5 (tools gate, 0 parse failures over ≥100 calls, ≥80 % of reference solve rate in the uncapped eval).
Outcome after review (GPT-6.1 Sol: AMEND) — accepted
Order: IQ2_M and Q2_K on 1 card → same files on 3 cards → IQ3_XXS on 1 card (tight resident candidate, not presumed paging) and 3 cards → IQ4_XS on 3 cards.
Fit is checked by measured placement at the stated context, per GPU; uneven splits watched.
Nothing is a finalist before the uncapped quality evaluation; deletion only of rejected candidates, after manifests.
MTP on/off is its own variable family; MTP use verified from engine logs and metrics.
Publish tool-call counts and reference solve counts next to every quality verdict.
0002-glm-quant-plan.review.md 0002-glm-quant-plan.review
AMEND. Keep the candidate quants, but correct the IQ3_XXS fit claim and make deletion depend on qualification. The strongest counter-argument is that weight capacity does not establish useful speed or quality : the spec already estimates one-card GLM at 8–13 tok/s, below the 15 tok/s registry floor.
IQ3_XXS is not automatically a paging experiment. Using the decision’s assumptions: 138.8 GB ≈ 129.3 GiB ; subtracting 21 GiB GPU residency leaves 108.3 GiB in RAM , about 6.7 GiB below the stated 115 GiB usable budget. Treat it as a tight resident-memory candidate. Declare paging only after accounting for context, runtime allocations and actual tensor placement.
All fit claims need a specified context and measured placement. Approximate host-weight budgets are 91.3 GiB for IQ2_M , 96.1 GiB for Q2_K , and 99.5 GiB for three-card IQ4_XS . Their remaining host margins are roughly 23.7, 18.9 and 15.5 GiB respectively. These exclude additional runtime memory. Likewise, allocating 66 GiB across three cards leaves only 2 GiB per card on average ; an uneven split can exhaust one GPU.
Budget bandwidth before promising 15 tok/s. Dual-channel DDR4-3200 has a theoretical maximum of 51.2 GB/s . At 15 tok/s, that permits at most 3.41 GB of DRAM traffic per generated token , with a smaller practical allowance. Measure effective bandwidth, expert placement and decode; VRAM capacity alone cannot show that the one-card route clears this constraint. Also, Q2_K is only 4.2% larger than IQ2_M—kernel efficiency and quality could easily determine the winner.
“~4-bit like accepted recipes” is precedent, not qualification. The spec’s reported sub-4-bit KLD of 0.28–0.38 makes quality a substantial risk. Establish the reference explicitly: exact EXL3 if loadable, otherwise the highest-bit GGUF that fits. If that is IQ4_XS, retain it through all candidate evaluations. A speed winner cannot be called a finalist before the uncapped quality evaluation.
Concrete alternative order: screen IQ2_M and Q2_K on one card, compare those files on three cards, then screen IQ3_XXS on both layouts and IQ4_XS on three cards. Run uncapped quality evaluation for promising configurations, then qualification for survivors. Give one-card IQ3_XXS an ordinary residency check before labelling it “stretch”; make three-card IQ3_XXS an explicit comparison point.
Replace “delete the losing 2-bit” with “delete a rejected candidate.” All four listed quants total 562.8 GB , leaving about 185 GB of the specified 748 GB free before Qwen weights, images and other artifacts. Build the combined inventory first. Preserve hashes and file references before deletion, and retain qualifying finalists under §9; being slower than another quant does not itself establish rejection.
Shared files do not establish shared MTP support. Verify that each pinned engine loads this exact revision and actually uses the retained MTP layer. Compare MTP off/on as a separate variable family; record accepted tokens/s, acceptance rate, settings and quality. Keep locally built ik_llama.cpp results experiment-only until its image provenance becomes eligible.
The decision’s quality summary omits essential acceptance conditions. Require the rolling 60-second decode floor , three qualification runs, cold-cache reset, no OOM/crash/repetition, and all six lab gates. Zero failures in 100 tool calls satisfies the spec but still gives an approximately 3% one-sided 95% upper bound on failure probability. The solve-rate test is also coarse: with four reference solves, all four are required; with five, four suffice. Publish these counts alongside occupied context and exact launch provenance.
0003-harness-interpretations.md 0003 — Harness interpretations of SPEC §5 (repetition, parse failures)0003 — Harness interpretations of SPEC §5 (repetition, parse failures)
Context: the T06a harness was exercised for 10 min against the operator's daily Qwen3.8-27B (EXL3 3 bpw, SGLang): 73 tool calls, 0 malformed tool calls, but 12 repetition hits — genuine reasoning loops (the same multi-paragraph block 3× inside one reasoning field) on sympy-21612 at 60–83k occupied tokens. Under §5 as written, the operator's everyday model fails stability.
Interpretations implemented by the harness:
Repetition, per field: the 64-token ×3 rule is evaluated within each generated field (reasoning, content, each tool call's arguments), not over the concatenation, because code drafted in reasoning and then emitted in a tool call trips the joined check without being a loop.
Parse failures, strict: besides invalid tool-call JSON and unparsed tool-call markup, every agent format error counts — including a response truncated at max_tokens (8192) with no tool call. Each response counts once; causes are split in detail.
Same-call streak (≥4 identical consecutive tool calls) resets at each task attempt.
Proposed (pending data, operator decides): keep §5 strict for this first pass and record all hits with causes; after T10/T20 produce real data, revisit whether "no repetition loops" should mean "no unrecovered loop" (a loop that ends the response by hitting max_tokens or ends the task) — not changed now.
Outcome after review (GPT-6.1 Sol: AMEND) — accepted
SPEC §5 amended: per-field repetition (cross-field diagnostic only), detector frozen for T10/T20, streak resets only on a fresh conversation (each task attempt starts a new mini-swe-agent conversation, so the harness already complies).
Three error categories: parse_failures (malformed tool calls only — the qualification criterion), output_truncations, agent_format_errors, all published.
No relaxation of "no repetition loops" now. Recovery and wasted-token figures to be published as supplementary data; any protocol revision later is prospective and does not re-grade earlier results.
Denominators always shown next to zero-failure claims.
0003-harness-interpretations.review.md 0003-harness-interpretations.review
AMEND. Keep the strict no-loop gate, but distinguish detector clarification from changes to eligibility. The strongest objection is that the harness simultaneously narrows repetition detection and broadens parse failures , while claiming to preserve §5.
Per-field repetition is defensible, but requires an explicit spec amendment. “Within one response” currently includes repetitions across fields. Adopt: “Within reasoning, content, or each tool-call argument field, independently, using the served model’s tokenizer.” Retain cross-field matches as diagnostics so the excluded cases remain auditable.
Do not label every agent-format error a tool-call parse failure. Invalid JSON and malformed tool markup qualify; an otherwise valid response exhausted at max_tokens without a tool call is a truncation or agent-protocol failure. Record tool_parse_failure, agent_format_error, and output_truncation separately. If all must disqualify, add that requirement explicitly to §5.
Resetting the same-call streak is appropriate only at an independent task attempt. Specify that a fresh agent conversation resets it; a retry continuing the same conversation does not. Otherwise attempt boundaries can conceal four consecutive identical calls.
The daily-model result does not justify relaxing the gate. Twelve genuine loops establish that this particular screening run fails the agreed criterion. They do not establish that the criterion is unsuitable for the two target models. The observed 60–83k occupancy also falls below the ≥128k requirement for “agentic-usable.”
Reject “unrecovered loops only” as the qualification alternative. Recovery does not erase wasted generation or latency. At exactly the specified decode floors, an 8,192-token response takes approximately 164 seconds for Flash-Next or 546 seconds for GLM , excluding prefill—27% or 91% of a ten-minute run. Publish recovery and wasted-token measurements as supplementary metrics while retaining zero detected loops for qualification.
The parse evidence is preliminary. Seventy-three calls falls short of the ≥100-call minimum. Under an independent binomial assumption, 0/73 failures still gives a one-sided 95% upper failure-rate bound of approximately 4.0% ; 0/100 gives 3.0% . Report the denominator without presenting zero observed failures as strong reliability proof.
Concrete alternative: ratify the field boundary and separate error categories before target comparisons; freeze that detector version for T10/T20; preserve each offending response and count affected responses separately from detector matches. Use later data to propose a prospective protocol revision, with existing results retaining their original qualification status.
0004-core2-quarantine.md 0004 — Quarantine CPU core 2 (logical CPUs 2 and 18)0004 — Quarantine CPU core 2 (logical CPUs 2 and 18)
Evidence: two hard resets on 2026-10-03 (04:57 near-idle, 06:56 during T10 model load). Each next boot logged an uncorrectable machine check in bank MC5 (Execution Unit) — CPU:18 MC5_STATUS[-|UE|MiscV|-|PCC|TCC|SyndV] and CPU:2 MC5_STATUS[-|UE|MiscV|AddrV|PCC|TCC|SyndV]. CPUs 2 and 18 are SMT siblings of physical core 2. No OOM, no GPU Xid, no kernel panic text: the core faulted and the platform reset. Likely cause: per-core Curve Optimizer/PBO undervolt too aggressive for that core, or core degradation (RMA case if BIOS is stock; BIOS 5101).
Decision (operator, option 2): take CPUs 2 and 18 offline for the whole campaign (resets at reboot; hygiene.sh check refuses to benchmark while they are online). Every record carries host.cpus_online/cpus_offline. llama.cpp and ik_llama launches use 15 threads. Then validate with tools/stability/cpu_check.sh (all-core AVX2, per-core single-thread boost bursts, all-core openssl) and look for new machine checks.
Consequence for results: numbers are for a 15-core 5950X. CPU-bound routes (GLM experts on CPU) are affected ~6 %; GPU/PCIe-bound trellis decode should not be. Any result produced before this decision does not exist (T10 never completed).
Check result (2026-10-03 07:25–07:40)
cpu_check.sh 300 20 180 on the 30 online CPUs: 0 machine checks, no reset (log: logs/cpu_check-core2-offline.log). Not conclusive for the near-idle failure mode; the hygiene check keeps the quarantine enforced for every run.
Revision (2026-10-03 10:30) — quarantine lifted with a tripwire
Operator BIOS check: PBO was Auto (ASUS stock); Curve Optimizer set to all cores, positive 0 (= no offset). So the CPU was effectively at stock when it faulted; a weak core 2 remains the leading hypothesis.
TARGET_CPUS="2 18" cpu_check.sh 300 20 180 with all 32 CPUs online: all-core AVX2, per-core bursts, 5 min sustained on CPU 2 and 18 each, 5 min on/off bursts on each, all-core openssl — 0 machine checks , no reset.
Not conclusive (both crashes came after hours of uptime). Core 2 is back online for full capacity; hygiene.sh check refuses to benchmark and re-offlines CPUs 2,18 if any machine check appears in the current boot; every record carries host.machine_checks_this_boot. llama.cpp/ik_llama launches back to 16 threads.
T10 ran on 15 cores (recorded in its host block). Trellis decode is PCIe-bound; one 16-core re-check is planned.
If it recurs at stock: keep core 2 offline for the campaign; RMA candidate.
Revision (2026-10-04 08:05) — permanent quarantine
Third fatal reset with core 2 online at stock: 2026-10-04 06:14, ~20.5 h into the boot, during the 16-core T10 re-check model load (CPU:2 MC5_STATUS[-|UE|MiscV|AddrV|PCC|TCC|SyndV], logged at the 06:54 boot). Same pattern as the 2026-10-03 06:56 crash (Flash-Next load) after a long uptime.
Core 2 stays offline for the rest of the campaign, enforced at boot by local-ai-bench-quarantine.service (tools/box/install.sh) and by hygiene.sh check (CPU_QUARANTINE defaults to "2 18"). All llama.cpp/ik_llama launches use 15 threads. The 16-core T10 re-check is dropped; all results are for a 15-core 5950X.
Stock settings + repeated MC5 UE on one core: AMD RMA candidate (operator's call).
0005-ik-llama-glm-tool-parsing.md 0005 — ik_llama.cpp does not parse GLM-5.3-Flash tool calls (commit 5bf8f0fe)0005 — ik_llama.cpp does not parse GLM-5.3-Flash tool calls (commit 5bf8f0fe)
Finding (T22b, IQ2_M, 1 card): the model emits well-formed GLM tool calls (<tool_call>bash<arg_key>command</arg_key><arg_value>…</arg_value></tool_call>) but ik_llama's server returns them as plain content with empty tool_calls: lab tools gate fails; stability run 0 tool calls / 47 parse failures (unparsed_markup). ik_llama builds tool parsers from the chat template (common/chat-auto-parser*); mainline llama.cpp b11312 parses the same GGUF correctly (tools gate passed in T20b). No matching ik_llama issue/PR found (search 2026-10-04).
Speed is unaffected and recorded: decode 10.0 tok/s (mainline 9.6), prefill 151 lab / 203 TTFT-8k (mainline ~80).
Decision: keep the ik_llama route as speed evidence only (experiment-only, fails §5 tools); do not patch ik_llama in this campaign — single-card GLM is ~10 tok/s on every route, below the 15 tok/s gate, so tool parsing would not change the outcome. Phase 3 option: upstream bug report with the captured response (artifacts/runs/…ik_llama…stability/proxy/parse_failures/).
3-card follow-up (2026-10-04, GPU2-first order, decision 0006)
No MTP, --fit-margin 4096: lab decode 13.4 / 14.0, agentic 13.3 tok/s steady (best GLM decode on this box), prefill 185–283; tool calls still unparsed (68 agent-format errors), tools gate fails.
MTP, --fit-margin 7168 (4096 OOMed creating the MTP context): lab decode 10.2 (acceptance 51 %), agentic 15.6 whole-run before the server crashed mid-run. The extra margin for the draft context pushes experts to host RAM.
Route parked: experiment-only speed evidence. The most interesting number for 0xSero is 13–14 tok/s without MTP; a tool-parser fix in ik_llama plus an attested image would be needed before it could be a 3-card candidate.
0006-pcie-layout.md 0006 — PCIe lane layout: what can be reconfigured, and what it buys0006 — PCIe lane layout: what can be reconfigured, and what it buys
Operator question (2026-10-04): can the PCIe lanes be reconfigured to get better model performance?
Measured topology (lspci, nvidia-smi, T05a)
Device Attach Link Measured H2D GPU0 3090 04:00.0 X570 chipset, behind a Gen4 x4 uplink shared with the KC3000 1 TB NVMe (root + all model files), NIC, BMC, SATA, USB Gen4 x8 to the chipset 6.0 GB/s GPU1 3090 0b:00.0 CPU lanes, slot 2 (x8/x8 split) — drives the display Gen4 x8 not measured (display) GPU2 3090 0c:00.0 CPU lanes, slot 1 (x8/x8 split) Gen4 x8 13.5 GB/s 990 EVO Plus 4 TB CPU lanes (M.2_1, Gen4 x4) unused: holds another OS install
Already optimal and not a lever: every link trains at Gen4 (16 GT/s); Resizable BAR is on (32 GB BAR1); ASPM is disabled on every GPU link. AM4 has 24 usable CPU lanes (16 slots + 4 M.2 + 4 chipset uplink); bifurcation or risers cannot add lanes, so the only real choices are which card sits where and how many cards share the 16 slot lanes .
Where PCIe bandwidth limits each route (evidence from our runs)
Flash-Next trellis (experts streamed from host RAM to one GPU): decode is PCIe-bound. 0xSero's data: ~26 GB/s → 62.7 tok/s, 14–16 GB/s → 53.2. Ours at 13.5 GB/s: lab 49–55, kit C1 46.9, agentic 43. x16 should add ~15–25 %.
llama.cpp/ik_llama prefill with experts in RAM is PCIe-bound (CPU-resident weights are copied to a GPU for large batches). Evidence: on GPU2, raising -b/-ub 512 → 2048 lifted GLM IQ2_M prefill 85 → 200 tok/s (each copy serves 4× more tokens). And 3 cards prefilled slower than 1 card (158 vs 200), consistent with the copy going to CUDA0 = GPU0, the 6 GB/s chipset card.
llama.cpp decode with experts on the CPU is DDR4-bound, not PCIe-bound (GLM ~9.5 tok/s on 1 card, 12.3 on 3 cards, from less RAM-resident weight). PCIe changes nothing here.
Options
# Change Cost Expected effect S1 3-card llama.cpp with the CPU-attached GPU2 as CUDA0 (CUDA_VISIBLE_DEVICES=2,1,0 in the container) none 3-card prefill back to ≥ 1-card level (~200+), no decode change S2 Larger ubatch (-b/-ub 4096) on llama.cpp routes none (more VRAM for compute buffers) prefill up to ~1.5–2× again; watch VRAM (the IQ2_M 3-card run already OOMed on CUDA0) P1 One card alone in slot 1 at x16 (T12, already planned) operator moves cards Flash-Next decode +15–25 % (the only route to ≥ 50 headline at full context); llama.cpp prefill ~2× vs x8; GLM exact mode at the bandwidth its 12.6 tok/s reference used P2 Plug the monitor into GPU0 (chipset card) instead of GPU1 move one cable GPU1 + GPU2 (both CPU x8) free for 2-card runs without going headless; desktop stays usable during single/2-card work P3 Model files on the CPU-attached 990 EVO Plus instead of the chipset KC3000 needs the operator's OK (other OS lives there) removes model loads and Flash-Next n-gram NVMe reads from the chipset uplink that GPU0 also uses; small effect except for 3-card runs with GPU0 P4 2 cards on CPU lanes (x8/x8), third card removed or idle none (GPU1+GPU2) avoids the 6 GB/s card; less VRAM (48 GB) so more experts in RAM
Decision (proposed)
Run S1 and S2 now in the current layout (software only, attributable, one variable family each), on GLM IQ2_M: 3-card GPU2-first at ub2048 (vs the existing GPU0-first record), then 1-card GPU2 at ub4096 (vs ub2048).
Keep P1 (T12) as the main hardware lever; in the x16 session also rerun the best GLM llama.cpp 1-card config so the x8→x16 effect on prefill is measured directly.
Recommend P2 to the operator at the T12 visit (same visit, one cable). P3 only if the operator wants to free the 990 EVO Plus; not worth touching another OS for this campaign.
Publish this as a site methodology section with the before/after records.
Outcome after review (GPT-6.1 Sol: AMEND) — accepted, plus operator rule
Bandwidth effects are hypotheses until measured. Expected-gain numbers above are estimates, not claims; x16 is the principal hardware hypothesis for Flash-Next, not a guarantee of qualification.
S1 is a fresh paired comparison, both headless, identical except card order: IQ2_M 3-card ub2048 with a 2 GiB --fit-target margin (the default-margin run OOMed on CUDA0 after 62 min), GPU0-first vs GPU2-first (CUDA_VISIBLE_DEVICES=2,1,0; the container sees the cards in PCI order). Every TTFT step now logs per-GPU PCIe rx/tx (nvidia-smi dmon -s t) to show which card receives the weight copies. Decode changes are reported, not assumed.
S2 : 1-card GPU2 ub2048 vs ub4096, both at 15 threads (the old ub2048 record ran 16 threads), peak VRAM and OOMs recorded.
P1 : rerun the best 1-card GLM llama.cpp config at x16 with the same bandwidth procedure (x8 control exists).
Operator rule (2026-10-04): GPU0 is used last. Single-card runs use GPU2; 2-card runs GPU2+GPU1; 3-card runs order the cards GPU2, GPU1, GPU0. Pending S1 evidence, all new multi-card launches use the GPU2-first order.
P2 offered to the operator at the T12 visit; P3 deferred (would need a §9 exception and touches another OS).
S1 result (2026-10-04 12:15–13:05 AEST) — card order, GLM IQ2_M 3-card ub2048, 2 GiB fit margin, headless
GPU0 first (control) GPU2 first (CUDA_VISIBLE_DEVICES=2,1,0) change TTFT prefill 8k / 32k 159 / 169 tok/s 240 / 242 tok/s +51 % / +43 % lab prefill 153 tok/s 218 tok/s +43 % lab decode C1 12.2 tok/s 11.8 tok/s −3 % PCIe rx during TTFT (avg / peak) GPU0 3.9 / 6.5 GB/s (link-saturated), others ~0.3 GPU2 5.7 / 14.3 GB/s, others ~0.5
The host→device weight copies for prefill go to CUDA0; putting a CPU-lane card there lifts prefill ~1.5×. Decode is within noise of unchanged (slightly lower). The GPU2-first config also finished a clean 10-min agentic run (11.3 tok/s, 24 tool calls, 0 parse failures, 0 loops) and passed the post-run gates without the OOM that hit the default margin. Rule confirmed: GPU0 last.
S2 result (2026-10-04 13:05–14:23 AEST) — ubatch, GLM IQ2_M 1 card GPU2, 15 threads
ub2048 (control) ub4096 change lab prefill 201 tok/s 288 tok/s +43 % TTFT prefill 8k / 32k 213 / 222 321 / 316 +51 % / +42 % lab decode C1 9.5 9.6 unchanged VRAM at ready 21.9 GB 22.9 GB +1 GB, no OOM in a 10-min agentic run
The 15-thread control matches the earlier 16-thread ub2048 record (9.5 / 200), so the core-2 quarantine costs nothing measurable on this route. Gain > 3 %, so the search continues: ub8192 on 1 card, and ub4096 on the GPU2-first 3-card config.
S2 step 2 (2026-10-04 16:13–16:30 AEST) — ub8192, 1 card GPU2
lab prefill 370 (+28 % vs ub4096), TTFT 384 / 404, but lab decode 9.1 (−5 %): the larger compute buffer makes --fit keep more experts in host RAM (VRAM at ready 22.2 GB). Prefill and decode now trade off, so the batching search stops here: ub4096 is the balanced setting (decode unchanged, prefill +43–51 %); ub8192 is the prefill-maximising variant, published as such.
3-card GPU2-first, ub2048 → ub4096 (16:30–16:58): lab prefill 218 → 258 (+18 %), TTFT 240 → 298 (+24 %), lab decode 11.8 → 11.2 (−5 %). Same prefill/decode trade-off as 1 card at ub8192; both published, no further steps.
0007-glm-current-layout-finalists.md 0007 — GLM current-layout finalists (T25a) and where the quality evaluation runs0007 — GLM current-layout finalists (T25a) and where the quality evaluation runs
Facts (all current-layout GLM records, 2026-10-03/04)
Route Cards Lab decode Lab prefill 10-min run (tool calls / parse failures / loops / solved) llama.cpp IQ2_M ub4096 1 9.6 288 27 / 0 / 1 / 1 of 2 llama.cpp IQ2_M ub2048, GPU2 first 3 11.8 218 24 / 0 / 0 / 1 of 2 llama.cpp IQ2_M ub2048, GPU0 first 3 12.3 158 26 / 0 / 0 / 1 of 2 (server later OOMed) llama.cpp IQ3_XXS ub2048 3 10.4 132 22 / 0 / 1 / 1 of 2 llama.cpp IQ4_XS ub2048, GPU2 first 3 8.3 156 25 / 0 / 0 / 1 of 2 llama.cpp Q2_K / IQ3_XXS 1 9.2 / 8.3 78 / 65 0–1 solved ik_llama IQ2_M (± MTP) 1 8.2–10.0 (MTP agentic ~12.6) ~150 0 tool calls: GLM tool calls not parsed (0005)
Every route is below the 15 tok/s registry gate on 1 card and on 3 cards (best 12.3). Decode is DDR4-bound: more cards only help by holding more experts in VRAM. Pending in the queue: IQ2_M 1-card ub8192 and 3-card ub4096 (prefill-only), ik_llama 3-card ± MTP (speed evidence only).
Proposal
No current-layout GLM finalist. §5 qualification requires decode ≥ 15 tok/s in every window, so the 3 × 10-min qualification runs would fail by construction. T25a records "no finalist: below the floor" for every current-layout config instead of running them. All configs stay published as documented experiments (§6 Phase 2).
One uncapped quality evaluation, of the best mainline weights, in the x16 session. Solve rate depends on weights, engine and sampling, not on the PCIe layout; the layout only changes how long the evaluation takes. So the T07 candidate evaluation of llama.cpp IQ2_M runs on the single x16 card next to the T23 reference (exact mode, or the highest-bit GGUF that fits if exact mode cannot load), with the identical task list, sampling and 60k-token budget. Each evaluation is estimated at 4–7 h (measured ~3–4k completion tokens per 10 min).
No uncapped evaluation for IQ3_XXS / IQ4_XS / Q2_K / ik_llama. They are slower than IQ2_M and cannot become candidates; their 10-min screening solve counts are published labelled "screening only". IQ4_XS only fits on 3 cards, so measuring it would cost ~6 h of current-layout time for an informational number.
Consequence: T12 (card move) can happen as soon as the current queue finishes (~3.5 h), not after ~1 day of current-layout quality runs. After the queue, IQ3_XXS and IQ4_XS are marked rejected and deleted (manifests kept) to make room for the x16 weights (Flash-Next 2.05/4.05 bpw: 63/108 GB; GLM exact EXL3: 125 GB).
Outcome after review (GPT-6.1 Sol: AMEND) — accepted
Below-floor is measured, not inferred from lab C1: every current-layout llama.cpp config with a 10-min run has a measured rolling 60-s generation-only minimum below 15 (IQ2_M 3-card 11.0–11.4, IQ2_M 1-card 8.8–9.1, IQ3_XXS 3-card 9.9, IQ4_XS 3-card 7.8). Configs without a stability run (prefill-only S1/S2 controls, ub8192) are recorded as "qualification skipped: subfloor lab C1; rolling floor not measured". T25a closes with no current-layout finalist once the pending queue (ik_llama 3-card ± MTP, ub follow-ups) is in; ik_llama stays experiment-only (0005).
Quality evaluation is conditional. At x16: measure IQ2_M 1-card first; only a config that qualifies (§5 incl. the rolling floor, ≥ 100 cumulative tool calls, post-run gates) gets the uncapped T07 candidate evaluation. An optional quality study of a non-qualifying config is labelled informational. Quality results apply to the tested configuration; layout independence is not assumed.
Runtime estimate corrected: the 60k budget is per task; worst case 2.5–3.3 h per task, 12–33 h per evaluation (less when tasks submit early). Budget accordingly; this is another reason to evaluate qualifiers only.
Labels: IQ4_XS, IQ3_XXS (3-card), Q2_K: "rejected for speed, quality unmeasured". IQ2_M 3-card GPU0-first: OOM after 62 min (default fit margin). IQ2_M/IQ3_XXS repetition hits are published with the run.
Storage: keep IQ2_M (x16 rerun) and IQ3_XXS (highest-bit GGUF that fits 1 card = fallback reference if exact mode cannot load); delete IQ4_XS after the queue, manifests and hashes kept for reacquisition.
0008-flashnext-3card-llamacpp.md 0008 — Flash-Next on 3 cards with llama.cpp: tune it as a separate 3-card candidate0008 — Flash-Next on 3 cards with llama.cpp: tune it as a separate 3-card candidate
Evidence (T11, 2026-10-04 15:38 AEST, IQ4_XS, GPU2,GPU1,GPU0 order, 2 GiB fit margin, -b 2048 / -ub 512 defaults)
Lab: all six gates pass , decode C1 57.1 / 56.9 tok/s (before / after the agentic run), lab prefill 608 / 605.
TTFT prefill: 987 tok/s at 8k, 852 at 32k.
10-min agentic run: 70 tool calls, 0 parse failures, 0 loops, 1 of 2 solved, whole-run decode 45.4, lowest 60-s window 38.6 (fails the 50 floor; decode falls as the context grows).
VRAM at ready 65.5 GB of 72; weights 97.7 GB, so ~30+ GB (incl. the n-gram embedding table) stays in host RAM.
For comparison, the 1-card trellis/SGLang baseline at x8: lab 49–55, kit C1 46.9, agentic 43.
Per SPEC §6 a 3-card config is a separate, clearly labelled recipe; llama.cpp routes are judged by the registry lab and §5 (the kit's SGLang-only panel cannot score a GGUF, so it is "kit band not measurable").
Proposal (current layout, before T12; one variable family per step, stop rule 3 × < 3 %)
Re-download IQ4_XS (deleted at 16:13 per plan; manifest kept, 97.7 GB, ~15 min).
Prefill : -b/-ub 2048, then 4096 (GLM showed +135 % then +43 %; target lab prefill > 1000).
Decode floor : the floor fails late in the run as context grows. One quant step down to IQ3_M (93.2 GB) or IQ3_XXS (88.0 GB) puts more experts in VRAM; run the better prefill setting from step 2 on it. Quality cost is unmeasured by the kit for GGUF; report solve counts and parse failures, label "quality: screening only".
If a config reaches lab decode ≥ 50, lab prefill > 1000 and holds the 50 floor in a 10-min run, it becomes a 3-card finalist: §5 qualification (3 runs incl. cold and max context, ≥ 100 tool calls) before T12.
Disk: the x16 downloads (296 GB) wait until this route is done and its rejected GGUFs are deleted.
Outcome after review (GPT-6.1 Sol: AMEND) — accepted
Route is experiment-only unless an accepted scoring path establishes the Flash-Next kit band for a GGUF; kit band recorded as unavailable, registry/headline eligibility unestablished. The T11 config is a failed screening config (floor −29.5 %, lab prefill −65 % vs targets), reported per constraint.
Batching steps, predeclared, judged by official lab prefill (decode and VRAM regressions checked): A -b 2048 -ub 2048, B -b 4096 -ub 4096, both vs the existing -b 2048 -ub 512 record. (GLM evidence says the ubatch sets the host→GPU copy amortisation, so b and ub move together.)
Before any quant change: decode vs occupied context from the existing T11 agentic trace (per-request generation speed against prompt tokens). IQ3_M is screened only if that shows a residency/transfer limit, not attention cost; its step is judged by the minimum rolling-window decode at matched occupancy. Expert/n-gram placement and memory use recorded from the server log.
Qualification only for a speed-passing config, naming its maximum occupied context (≥ 4k headroom), cold resets, post-run gates, ≥ 100 tool calls cumulative. Phase timebox kept; x16 is not deferred beyond this route's steps.
Decode vs occupied context (T11 3-card agentic trace, 55 requests, generation-only tok/s per request)
prompt tokens per-request decode 1.5k–6k (django task) 35–51, typical ~42 6k–20k 30–49, typical ~40 20k–31k 30–41, typical ~38 31k–39k 28–38, typical ~33
Two separate gaps to the 50 floor: a base gap (~42 tok/s already at 2–6k tokens, vs lab C1 57 measured over one long generation; agentic requests are short, 70–600 tokens) and a context slope of roughly −20 % from 5k to 39k. One quant step (IQ3_M, −4.6 % of bytes; IQ3_XXS −9.9 %) cannot close a 25–30 % gap even if every saved byte moved experts to VRAM. IQ3_M step dropped ; the predeclared batching steps still run (prefill evidence, cheap). After them the route closes as experiment-only and its GGUF is deleted for the x16 weights.
Batching result (2026-10-04 18:00–18:23 AEST) — route closed
-b / -ub lab decode C1 lab prefill TTFT 8k / 32k 2048 / 512 (T11 record) 57.1 608 987 / 852 2048 / 2048 44.3 (−22 %) 717 (+18 %) 1067 / 931 4096 / 4096 35.0 (−39 %) 713 980 / 881
Larger ubatches buy little prefill and cost a lot of decode: the bigger compute buffers make --fit move experts to host RAM. No setting reaches lab prefill > 1000 while keeping decode ≥ 50, and the agentic floor already failed at the default. Route closed: experiment-only , best config = the default-batch T11 record (all six gates, lab 57.1, agentic floor 38.6). GGUF deleted for the x16 weights (manifest kept).
0009-no-x16-session.md 0009 — No x16 session: Phase 1 and T23 continue in the current layout (operator decision)0009 — No x16 session: Phase 1 and T23 continue in the current layout (operator decision)
Operator, 2026-10-04 19:15 AEST: skip the x16 test; the expected difference is judged marginal. T12 (card move) and T05b are cancelled. The operator's call stands; the estimate it overrides is recorded for the reader: 0xSero's reference gives +17.9 % decode for 14–16 → ~26 GB/s, and our x8 baseline (lab C1 49–55, kit C1 46.9, agentic 43, floor not held) sits just under the 50 tok/s target, so reaching it now depends on tuning at x8.
Re-plan (single card = GPU2, CPU-lane x8, 13.5 GB/s; GPU1 measured identical)
T13 PR #144 rerun at x8 against image f77c72f6 (lab, kit sweep with early exit off, TTFT, panel, 10-min run). Proof for 0xSero is "x8 on this box"; whether he accepts it is asked in Phase 3.
T15 brought forward: the main recipe on EXL3 2.05 and 4.05 bpw, only the weights changed. At x8 the expert bytes per token are the bottleneck, so 2.05 bpw is the most likely route to ≥ 50 sustained; the kit panel decides whether it stays in the quality band (required for the headline and registry PR). 4.05 bpw is the quality anchor.
T23 GLM exact mode on GPU2 at x8 (route test + quality reference; may not load: ~119 GB pinned on ~125 GiB).
T14a/T14b grid on the best quant from 1–2 (expert cache GB × context × KV format; then prefill chunk and n-gram tier), designed as its own decision once 1–3 are in. T16 qualification and headline follow on the winner.
GLM IQ2_M x16 rerun dropped; IQ3_XXS kept as the fallback reference.
Outcome after review (GPT-6.1 Sol: AMEND) — accepted
Recorded as a deliberate restriction of the search, not evidence that x16 is marginal. Gap to target at x8: kit C1 46.9 → 50 needs +6.6 %; agentic 43 → 50 needs +16.3 %; there is currently no qualifying Flash-Next result.
Order: T13 at 3.05 → matched 2.05 / 4.05 screens (same launch, same context) → kit panel as a filter (top-1 ≥ 0.987, KL ≤ 0.0013) before any headline time → bounded tuning of in-band candidates, 3.05 included . Out-of-band results are published with kit eligibility, never as the registry path. "Bytes per token" is a hypothesis; transfer/cache evidence is recorded where the engine exposes it.
Acceptance unchanged: every eligible rolling 60-s window ≥ 50 and lab prefill > 1000 on the same config at the claimed occupied context. If no x8 config meets speed and quality together, that outcome is published as is.
T23 at x8: exact mode may be the GLM quality reference even below 15 tok/s (not a recipe); peak host memory and any load failure recorded (~110.8 GiB pinned vs ~125 GiB MemTotal). If it cannot load, the reference is the highest-bit GGUF that actually fits one card here, determined by a load test, not assumed to be IQ3_XXS.
Timebox and the 3 × < 3 % stop rule stay. Cancelled tickets (T12, T05b, x16 parts of T13/T14/T23) are recorded as skipped with the operator's reason.
0010-cpu-boost-off.md 0010 — CPU boost off for the rest of the campaign0010 — CPU boost off for the rest of the campaign
Evidence (2026-10-04 19:48 AEST): hard reset during the T13 Flash-Next SGLang model load with core 2 already offline. The next boot decoded CPU:1 MC5_STATUS[-|UE|MiscV|AddrV|PCC|TCC|SyndV], "uncorrected error caused a data fabric sync flood" — the same bank, status and syndrome as the core-2 faults (0004), now on core 1 . Tally of fatal resets: 5, of which 3 during a Flash-Next SGLang load (expert repacking: heavy multi-threaded host work) and 2 after long uptime. Temperatures at idle are normal (Tctl ~56–60 °C). A second core failing the same way points to the CPU (or its voltage under boost transients) rather than one bad core.
Decision: disable Precision Boost from Linux (/sys/devices/system/cpu/cpufreq/boost = 0, all cores capped at the 3.4 GHz base clock, lower voltage), applied at boot by local-ai-bench-quarantine.service and immediately on 2026-10-04 20:12. Core 2 stays offline. Every record now carries host.cpu_boost and host.cpu_max_mhz.
Consequence: CPU-bound work (GLM experts on CPU, model loads) gets slower; PCIe-bound trellis decode should not. Records before 20:12 ran with boost on and are labelled by that field's absence (= boost on). If faults continue with boost off, the CPU is not usable for unattended runs and the remaining work pauses for the operator (RMA).
0011-flashnext-x8-tuning-grid.md 0011 — Flash-Next tuning at x8 (T14): grow the expert cache0011 — Flash-Next tuning at x8 (T14): grow the expert cache
Evidence
T15 is not runnable: the offload engine's HostExpertStore only supports 3-bit experts (offload_store.py: if bits != 3: raise NotImplementedError), so 2.05 and 4.05 bpw crash at load in both images. Only 3.05 bpw remains (the CPU-expert fallback, SGLANG_EXL3_MOE_OFFLOAD=cpu, is far slower and not pursued).
T13 (PR #144 launch, f77c72f6, x8, boost off): lab C1 54.5–55.6, lab prefill 2,433–2,596, kit C1 48.2 (32k: 49.9), kit panel in band (top-1 0.9887, KL 0.00097), agentic whole-run 44.1, min 60-s window 37.9 (floor 50 fails).
Agentic per-request decode is flat across context (median 43.7 / 43.8 / 41.7 tok/s at <8k / 8–20k / 20–40k), so context length is not the limiter; the gap to lab C1 (55) is consistent with expert-cache misses on coding traffic (each miss streams experts over x8 PCIe). Hypothesis, to be tested by changing the cache size.
VRAM budget at ready (server log): expert cache 11.08 GB (auto-fit: free 19.82 − staging 0.48 − reserve 8.26), KV pool 5-bit 1.77 GB for 210k tokens, 3.23 GB still free after CUDA graphs . Context length is therefore a weak lever (the whole KV pool is 1.77 GB); the reserve is the strong one.
Grid (base = PR #144 launch on the latest main-built image a21efea6, GPU2 x8; one variable per step)
Step Change Expected expert cache R1 SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB 8.26 → 6.0~13.3 GB (+20 %) R2 reserve → 5.0 (only if R1 is stable and gains ≥ 3 %) ~14.3 GB (+29 %) N1 SGLANG_EXL3_NGRAM_TIER nvme → pinned (32.6 GB pinned RAM) on the best Rn-gram reads from RAM G1 --max-running-requests 1 --cuda-graph-max-bs-decode 1 (single-stream agentic) on the best, then re-fit the reservefrees graph memory
Metric per step: minimum rolling 60-s window and median per-request decode in the 10-min agentic run (the floor), with lab C1 and lab prefill checked for regressions. Every step runs lab,ttft,stab; the kit panel is re-run on the winner only (cache size does not change the math). Stop rule: 3 consecutive steps < 3 % on the floor metric. If no step lifts the minimum window to ≥ 50, Phase 1 closes with "not reached at x8" , publishing the best config.
Outcome after review (GPT-6.1 Sol: AMEND) — accepted
Cache misses are a hypothesis: the floor needs +31.9 % (26.4 → 20.0 ms/token); a 20–29 % larger cache need not deliver that. Expert-cache hit/miss and transfer stats are recorded if the engine logs them, else marked unavailable.
Order: R0 (unchanged PR #144 launch on a21efea6) → R1 reserve 6.0 → G1 single-stream graphs on R1 → reserve refit only if peak VRAM (prefill and max-occupancy) shows room → N1 pinned n-grams with host-RAM headroom recorded. R2 (5.0) is not run blind: measured headroom after R1 would be ~0.97 GB, after R2 ~−0.03 GB.
max-running-requests 1 is a recipe concurrency constraint, stated as such.
Floor = minimum rolling 60-s generation-only window after the first 60 s (harness definition); window coverage reported. One 10-min run screens; the winner then needs full §5 qualification, kit band, and ≥ 128k occupied for "agentic-usable". Closure wording: "50 tok/s floor not reached in the tested x8 configurations".
R0 result and re-plan (2026-10-04 22:42 AEST)
R0 (a21efea6, PR #144 launch unchanged): load 260 s, lab C1 55.8, lab prefill 2,419, kit C1 47.4 (32k 48.6), TTFT 2,009 / 2,681, panel in band (top-1 0.9897, KL 0.00097) — then the 10-min agentic run crashed: CUDA OOM . After serving starts, Triton kernels load lazily (log: "device-loaded after serving started … free 0.83 → 0.09 GiB") and consume the 3.23 GB that looked free at ready. So the 8.26 GB reserve is needed on this image; shrinking it cannot grow the expert cache safely. (f77c72f6 survived the same agentic run at the same reserve.) Re-plan: R1 (6.0) is already running and is kept as evidence; R1+G1 dropped . Next: R0 repeat (is the OOM reproducible?), G1 at reserve 8.26 (single-stream frees graph and request memory → headroom), N1 pinned n-grams.
R1 (reserve 6.0, 22:42): load OK, then CUDA OOM at the first lab request — reserve cannot go below 8.26.
R0 repeat, agentic run only (22:58–23:22): no OOM ; 71 tool calls, 0 parse failures, whole-run 43.3, min window 39.7; lab after 57.4 / 2,537. The first R0's OOM is intermittent (it followed lab, kit sweep, TTFT and panel on the same server). Reported as an a21efea6 stability risk for PR #144; frequency unknown from 2 runs.
G1 single-stream at reserve 8.26 (23:22): min window 40.1 (+1 % vs R0 repeat), whole-run 43.1, lab 54.6–55.6.
N1 pinned n-grams (23:48): min window 40.9 (+3.1 %), whole-run 44.6, lab C1 53.7 / 51.3 (lower), prefill unchanged.
Closure (2026-10-05 00:11 AEST)
Stop rule met: R1 no gain (OOM), G1 +1 %, N1 +3 % (within run-to-run spread: two R0 runs and T13 span 37.9–39.7). Phase 1 result: the 50 tok/s floor was not reached in the tested x8 configurations. Best measured agentic floor ~40–41 tok/s (−18 to −20 %); lab C1 55–57, lab prefill ~2,400–2,600, kit panel in band. Best config for publication: the PR #144 launch (image a21efea6, or f77c72f6 which had no OOM in its run). No T16 finalist; T18 registry recipe not prepared; the PR #144 rerun evidence and the kit submission (with this tuning table) are prepared locally for Phase 3.
0012-glm-outcome.md 0012 — GLM-5.3-Flash outcome (T24, T25a, T25b, T26)0012 — GLM-5.3-Flash outcome (T24, T25a, T25b, T26)
Result: no GLM-5.3-Flash configuration on this box reaches the registry's 15 tok/s gate; every route is a documented experiment, and no GLM registry PR is prepared. All numbers: lab C1 decode / lab prefill, current layout.
Route Cards Decode Prefill Tools Classification llama.cpp b11312 IQ2_M, ub4096 1 9.6 288 pass, 0 parse failures experiment (1 card < 15) llama.cpp IQ2_M, ub2048, GPU2-first 3 11.8 218 pass experiment (3-card, < 15) llama.cpp IQ3_XXS / IQ4_XS 3 10.4 / 8.3 132 / 156 pass rejected for speed, quality unmeasured llama.cpp Q2_K / IQ3_XXS 1 9.2 / 8.3 78 / 65 pass rejected for speed ik_llama 5bf8f0fe IQ2_M (no MTP) 3 13.4–14.0 185 fails (tool calls unparsed, 0005)experiment-only, unattested image ik_llama IQ2_M + MTP 1 / 3 8.2–12.6 ~150–166 fails experiment-only; OOM / server crash 0xSero glm53-flash-offload, exact mode (x8) 1 7.4–7.5 720–739 pass, 0 parse failures quality reference; route below gate at x8
T25a: no current-layout finalist (decision 0007; rolling-floor failures measured in every 10-min run).
T24: the uncapped candidate-vs-reference comparison runs only for finalists, so it is not run. Exact mode loads on this box (~120 GB host RAM, 23.3 GB VRAM) and stays available as the reference if a future config qualifies.
T25b: 1-card routes < 15 → experiment; 3-card routes < 15 → no separate recipe; ik_llama → experiment-only until a tool-parser fix and an attested local-ai-images build (operator to raise with 0xSero).
T26: not applicable.
Hardware context for the reader: dual-channel DDR4-3200 (~16.6 GB/s host memcpy measured), x8 PCIe (13.5 GB/s), 15 cores at base clock (decisions 0004, 0010). Fast mode (0xSero's 28.9 tok/s route) needs ~238 GB RAM: not possible on AM4.
Budget and stop-rule log (budget.md) Budget
Phase Started Budget Status 0 setup 2026-10-02 1 day in progress 1 Flash-Next – 2–3 days – 2 GLM – 2–3 days – 3 review & PRs – operator-paced –
Stop-rule log
Date Route Last 3 gains Decision
Site rules Status badge not started = no records; done = a config with a kit / registry / kit+registry / experiment-only eligibility whose stability verdict is qualified; otherwise running. site/status.json overrides it (shown as override). Stability verdict failed = any stability run failed or crashed, or shows parse failures, repetition hits or crash/OOM; qualified = at least 3 passing runs, one of them cold, and at least 100 tool calls across them; screened = passing runs short of that. Flash-Next target one passing lab record of the config with decode >= 50 tok/s and lab prefill > 1000 tok/s. Headline = the qualified target-meeting config with the largest occupied context reached in its passing stability runs. GLM target best lab C1 decode of a 1-card config against the registry speed gate (pins.json). 3-card configs are shown separately and do not satisfy the 1x RTX 3090 request. Agentic-usable qualified and a passing stability run reached >= 128k occupied tokens; otherwise the label is 'configured for N k'. Kit band top-1 >= 0.987 and mean KL <= 0.0013 (Flash-Next, quality_panel records). GLM quality floor solve-rate ratio >= 80% of the reference (metric solve_rate_ratio on quality_eval records, else computed from the model's eligibility=reference quality_eval). Dropped config a config with a failed or crashed record and no record carrying a positive eligibility.
Pins 0xSero/local-ai-registry d21258dd744e7c78be28177c90060c6af8e10b7e0xSero/local-ai-recipe-kit ef883d269e50ecbea290f919f095e2f3ca633b42SWE-agent/mini-swe-agent {"commit": "04d809ceab9df28f9adaed044884180159172930", "nearest_tag": "v2.4.6"}panel qwen3.8-flash-next-exl3 {"path": "reference/qwen3.8-flash-next-exl3-ref-panel.json", "sha256": "3350eef4e980779256a0c1e872d9345d0d21bdf43690e1958526aef17e02488d"}panel glm-5.3-flash-exl3 {"source": "0xSero/local-ai-recipe-kit PR #3 @ b5d0d3c2d567c4d4c1a295acc0a0fc86470b9097", "path": "artifacts/glm-5.3-flash-exl3-ref-panel.json", "sha256": "3445baefa35851f7f431b4f46fc2266a9d234d79526b1d4fc1216093f552f269"}lab gates load, chat, reasoning, tools, context, speed lab speed_min_tps 15.0 lab speed C1 streamed decode over first 30 s after first token, temperature 0.8, uncapped answer lab context needle in a prompt of ~0.85*ctx; pass if recalled and prompt_tokens >= 0.6*ctx lab chat finish_reason=stop with non-empty content lab reasoning separate reasoning_content and 17*23=391 in content lab tools get_weather(city=Paris) call, then tool result (17) used in reply lab prefill_definition context-gate prompt_tokens / total request seconds (includes generation) kit early_exit_threshold 0.95 kit early_exit_policy disabled during exploration (spec §4)
pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78, pinned 2026-10-02
Protocol (from SPEC.md) 2. Hardware ("the ryzen", dedicated to this work)
Part Detail CPU Ryzen 9 5950X, 16C/32T, AVX2, no AVX-512 Board ASUS Pro WS X570-ACE RAM 128 GB DDR4-3200, dual-channel (4×32), ~125 GiB MemTotal GPU 3× RTX 3090 24 GB, driver 610.57.04 PCIe today GPU0 04:00.0 x8 behind the X570 chipset; GPU1 0b:00.0 x8 CPU (drives display); GPU2 0c:00.0 x8 CPU PCIe x16 layout one card alone in the primary CPU slot (operator moves cards) Storage 1 TB NVMe, ~748 GB free: active models. 20 TB ZFS HDD: unused
Identify GPUs by UUID in every result. Never benchmark on the display GPU. 3-card runs are headless : the operator logs out of Hyprland and everything is driven over SSH; each record notes display: off.
4. Definitions and metrics
Official speed : the numbers from lab/lab.py try --on endpoint in the registry. Decode = C1 over the first 30 s at T=0.8. Lab prefill = context-gate prompt tokens ÷ total request seconds.
Lab prefill is a conservative number. A value > 1000 counts as proof that prefill clears 1000. Record time-to-first-token prefill separately, with the prefix cache disabled and fresh prompts.
Pinned harness : at Phase 0 start, record the commit SHAs of the registry, kit, mini-swe-agent and the reference panels. Every run cites them, and the six lab gates and thresholds are those at that commit.
Decode floor measurement : generation-only tok/s (output tokens ÷ time spent streaming them, excluding prefill, tool execution and idle time) over rolling 60 s windows, ignoring the first 60 s. "Holds the floor" = every window ≥ the floor. This is separate from official lab C1 and from the whole-run average.
Supplementary speed : the kit PROTOCOL.md sweep (prefill 8k/16k/32k/64k, decode C1–C4), always run with early exit disabled while exploring.
Real-world speed : decode measured across the whole agentic run.
Context is reported as occupied tokens (the prompt actually sent), alongside the configured limit. Maximum tested: 260k. The label "agentic-usable" is awarded only after a config passes §5 qualification at ≥ 128k occupied tokens ; otherwise it shows "configured for N k". Shorter profiles are still published.
Occupied context = prompt tokens counted by the served model's own tokenizer (usage.prompt_tokens from the server), excluding generated tokens. A max-context claim needs evidence that a request actually reached that occupancy, with ≥ 4k tokens of generation headroom left.
MTP / speculative decoding : always report accepted tokens/s, acceptance rate and the draft settings.
5. Stability and quality protocol
Run types
Stability runs : capped at 10 min each (screening and qualification below).
Quality evaluation (GLM solve-rate, §5 GLM quality floor): a separate run with no 10-min cap but a fixed token budget per task. Finalists only.
Agentic harness
mini-swe-agent, pinned commit.
Fixed set of 5–10 SWE-bench-Lite tasks, at least one growing past 64k context.
Sampling is fixed: temperature, top-p, seed where supported, max tokens.
Screening run
Finalist qualification
3 runs of ≤ 10 min each : one starting cold, one at the claimed maximum occupied context, one repeat.
Cold = container restarted (prefix/KV and expert caches empty), the n-gram row cache empty, and the OS page cache dropped (sync; echo 3 > /proc/sys/vm/drop_caches). Each record lists which caches were reset.
Each run works through the task set in fixed order, looping, until the 10 min wall clock runs out. An unfinished task counts as unsolved.
Pass criteria (every run)
No crash and no OOM.
0 tool-call parse failures . The ≥ 100 tool-call minimum is counted across all runs of a config (screening + qualification + extra 10-min runs as needed), not per run.
No repetition loops. Detector (frozen per decision 0003): any 64-token sequence repeated ≥ 3 times within one field of a response (reasoning, content, or one tool call's arguments, each checked independently; cross-field matches are logged as diagnostics only), or the same tool call (name + arguments) issued ≥ 4 times in a row within one agent conversation (the streak resets only when a fresh conversation starts). When it triggers, save the offending response. Tokens come from the served model's tokenizer when the server exposes one, else tiktoken cl100k (recorded).
Error categories are reported separately : malformed tool calls (invalid argument JSON, unparsed tool-call markup, unknown tool, missing command) are tool-call parse failures ; responses cut off at max_tokens without a tool call are output truncations ; other agent-protocol errors are agent format errors . Only the first category is the "0 parse failures" criterion; the other two are published next to it.
Decode stays at or above the absolute floor for the whole run: Flash-Next ≥ 50 tok/s; GLM ≥ 15 tok/s when the config is a registry candidate.
All six registry lab gates still pass after the run.
GLM quality floor
Passes the registry tools gate.
Meets the parse-failure rule above.
Solves ≥ 80 % as many tasks as the reference on the same task set, measured in the uncapped quality evaluation with identical task IDs, sampling settings and per-task token budget for candidate and reference. If the reference solves 0 tasks, the solve-rate criterion is reported as not measurable and the config cannot be a registry candidate on quality grounds.
Reference: GLM exact mode (bit-exact EXL3 3.05 bpw). If exact mode cannot load, the reference is the highest-bit GGUF that fits.
Also report KLD and top-1 agreement where they are available.
Flash-Next quality
Kit band: top-1 ≥ 0.987 and KL ≤ 0.0013 against the exllamav3 reference panel (tools/score_ref_panel.py, the kit's pinned panel).
The band is required for the headline result and for the registry PR . Configs outside it are still submitted to the kit and published, labelled "outside band" (precedent: kit PR #1).
Generated 2026-10-05 06:51 AEST by tools/site/build.py from 152 records in results. Private preview.