Skip to content
← Blog

Mouse on GLM-5.3-Flash

23 of 30 FrontierHarness tasks on Z.ai's Flash model, next to the published GLM-5.3 harness runs, for $6.72 in tokens.

Community results shared this week ran GLM-5.3 and GLM-5.3-Flash from Z.ai through five coding agent harnesses on the 30 FrontierHarness tasks. Z.ai's own harness, ZCode, came out on top, and was the only one run on Flash. We ran Mouse on GLM-5.3-Flash on the same tasks. It passed 23 of 30, at $0.29 per pass.

Pass rate
77%
23 of 30 tasks
Cost per pass
$0.29
$6.72 for all 30
Cache hit rate
97%
of input tokens
Median time
4:11
per pass
MouseZCodeCodexClaude CodePiOpenCodeDSH CreatorDSH Standard / DSHDSH PTCDSH MinimalOh My PiKimi CodeExo HarnessHermes $0.2$0.5$1$2$5$10$2050%60%70%80% Mouse + Kimi K383.3% · $2.79 · RuntaMouse + GLM-5.3-Flash76.7% · $0.29 · RuntaZCode + GLM-5.382.2% · $1.98Claude Code + GLM-5.380.0% · $2.58Codex + GLM-5.366.7% · $1.19Pi + GLM-5.360.0% · $2.10OpenCode + GLM-5.346.7% · $2.28ZCode + GLM-5.3-Flash75.6% · $0.15DSH + DeepSeek V4.1 Flash73.3% · $0.21Codex66.7% · $3.47DSH Creator63.3% · ≥$3.29Claude Code63.3% · ≥$18.34Pi60.0% · $2.43DSH PTC60.0% · ≥$4.58DSH Standard60.0% · ≥$3.46Oh My Pi56.7% · $4.75Kimi Code56.7% · $3.65DSH Minimal56.7% · ≥$4.72Exo Harness53.3% · $1.04OpenCode50.0% · ≥$3.24Hermes50.0% · ≥$2.90 Total recorded cost / passes (log scale) Pass rate Names without a model are the published Kimi K3 baselines; the orange line is their cost frontier. ≥ marks incomplete baseline cost coverage. GLM and DeepSeek rows are third-party local runs at list prices on observed cache. Mouse rows are Runta runs under FrontierHarness's scripts.
FrontierHarness tasks, pass rate against total recorded cost divided by passes. Different models, providers, and runtimes: a side-by-side, not a controlled ranking.

The comparison

The published FrontierHarness board fixes the model at Kimi K3. The GLM results are a separate, community-run extension: the same 30 tasks and scoring, with GLM-5.3 on a Zhipu Max plan instead. Cost is observed tokens multiplied by Z.ai's API list price, not what the plan bills. ZCode was run three times per model, so its rows are averages.

Harness Model Pass rate Passed $ / pass Runs
ZCode GLM-5.3 82% 24.67/30 $1.98 26 / 22 / 26
Claude Code GLM-5.3 80% 24/30 $2.58 1
Mouse GLM-5.3-Flash 77% 23/30 $0.29 1
ZCode GLM-5.3-Flash 76% 22.67/30 $0.15 24 / 21 / 23
Codex GLM-5.3 67% 20/30 $1.19 1
Pi GLM-5.3 60% 18/30 $2.10 1
OpenCode GLM-5.3 47% 14/30 $2.28 1

On Flash, Mouse's single run (23) sits within ZCode's three runs (24, 21, and 23). ZCode is cheaper per pass, $0.16 against Mouse's $0.29. On the full GLM-5.3 model, ZCode and Claude Code pass 24 to 25 tasks and cost $2 to $2.60 per pass, so Mouse on Flash comes within one or two tasks of them at a seventh to a ninth of the cost per pass.

Against Kimi K3

The same harness scored 25 of 30 on Kimi K3 at $2.79 per pass. Flash lost two tasks that K3 passed, both DeepSWE: expr-try-catch-errors and fastapi-deprecation-response-headers. It did not pass anything K3 failed. The other five misses were shared: meriyah on DeepSWE, and dna-insert, polyglot-c-py, extract-elf, and gcode-to-text on Terminal-Bench.

The cost difference is mostly the model's price. Flash lists at $0.15 per million input tokens, $0.03 cached, and $0.50 output. Cache hits covered 96.7% of input tokens, so the median passing task cost a little over a cent. The DeepSWE tasks were most of the bill: the nine of them cost $6.10 of the $6.72.

Terminal-Bench

Task GLM-5.3-Flash Steps Cost Time Kimi K3
regex-log Pass 26 $0.05 8:56 Pass
openssl-selfsigned-cert Pass 26 $0.01 3:28 Pass
polyglot-c-py Fail 17 $0.04 9:42 Fail
sqlite-db-truncate Pass 12 $0.01 3:46 Pass
git-leak-recovery Pass 15 $0.01 3:20 Pass
log-summary-date-ranges Pass 18 $0.01 3:12 Pass
constraints-scheduling Pass 13 $0.01 3:56 Pass
gcode-to-text Fail 57 $0.13 16:26 Fail
dna-insert Fail 16 $0.02 6:48 Fail
largest-eigenval Pass 15 $0.04 16:48 Pass
merge-diff-arc-agi-task Pass 23 $0.02 5:26 Pass
vulnerable-secret Pass 16 $0.01 3:16 Pass
extract-elf Fail 21 $0.03 7:02 Fail
build-cython-ext Pass 98 $0.17 12:16 Pass
kv-store-grpc Pass 21 $0.01 3:20 Pass
chess-best-move Pass 17 $0.01 4:11 Pass
db-wal-recovery Pass 14 $0.01 4:27 Pass
code-from-image Pass 8 $0.00 2:13 Pass
modernize-scientific-stack Pass 12 $0.01 2:43 Pass
multi-source-data-merger Pass 12 $0.01 3:44 Pass
sanitize-git-repo Pass 15 $0.01 3:03 Pass

DeepSWE

Task GLM-5.3-Flash Steps Cost Time Kimi K3
anko-typed-variable-bindings Pass 114 $0.38 20:55 Pass
arktype-json-schema-refs-dependencies Pass 211 $0.82 48:27 Pass
fastapi-deprecation-response-headers Fail 161 $0.54 42:52 Pass
httpx-multipart-response-parsing Pass 96 $0.34 31:36 Pass
expr-try-catch-errors Fail 163 $0.76 42:22 Pass
python-statemachine-state-data-scoping Pass 250 $1.55 77:48 Pass
katex-multicolumn-array-spans Pass 127 $0.53 33:43 Pass
scc-bounded-memory-spilling Pass 100 $0.48 29:46 Pass
meriyah-explicit-resource-declarations Fail 154 $0.71 39:01 Fail

How the run was made

  • FrontierHarness's own trial driver, unmodified, at their eval commit e837a70, on Runta with a fresh restore for every task and 4 vCPU / 8 GiB per task. Terminal-Bench through Harbor 0.22.0, DeepSWE through Pier 0.3.1.
  • Harness commit 315e2b8 on OpenCode 1.18.27, the same checkpoint as the Kimi K3 run, with only the model changed: glm-5p3-flash served by Fireworks.
  • Cost is each trial's observed tokens at Z.ai's GLM-5.3-Flash list price, which Fireworks matches. Three attempts on two tasks hit Ubuntu mirror errors while setting up the container, before the model ran; they were re-run and are not counted as failures.
  • One run, 2026-09-14. The other rows were run locally by a third party, on a different provider, runtime, and harness versions (Claude Code 2.1.237, OpenCode 1.18.19, Codex 0.148.0, Pi 0.84.2), so this is a side-by-side, not a controlled ranking.

GLM results for ZCode, Claude Code, Codex, Pi, and OpenCode are from a community comparison shared this week, building on work by LotusDecoder. Mouse is built on OpenCode and is not affiliated with it or with Z.ai.