Mouse on GLM-5.3-Flash
23 of 30 FrontierHarness tasks on Z.ai's Flash model, next to the published GLM-5.3 harness runs, for $6.72 in tokens.
Community results shared this week ran GLM-5.3 and GLM-5.3-Flash from Z.ai through five coding agent harnesses on the 30 FrontierHarness tasks. Z.ai's own harness, ZCode, came out on top, and was the only one run on Flash. We ran Mouse on GLM-5.3-Flash on the same tasks. It passed 23 of 30, at $0.29 per pass.
The comparison
The published FrontierHarness board fixes the model at Kimi K3. The GLM results are a separate, community-run extension: the same 30 tasks and scoring, with GLM-5.3 on a Zhipu Max plan instead. Cost is observed tokens multiplied by Z.ai's API list price, not what the plan bills. ZCode was run three times per model, so its rows are averages.
| Harness | Model | Pass rate | Passed | $ / pass | Runs |
|---|---|---|---|---|---|
| ZCode | GLM-5.3 | 82% | 24.67/30 | $1.98 | 26 / 22 / 26 |
| Claude Code | GLM-5.3 | 80% | 24/30 | $2.58 | 1 |
| GLM-5.3-Flash | 77% | 23/30 | $0.29 | 1 | |
| ZCode | GLM-5.3-Flash | 76% | 22.67/30 | $0.15 | 24 / 21 / 23 |
| Codex | GLM-5.3 | 67% | 20/30 | $1.19 | 1 |
| Pi | GLM-5.3 | 60% | 18/30 | $2.10 | 1 |
| OpenCode | GLM-5.3 | 47% | 14/30 | $2.28 | 1 |
On Flash, Mouse's single run (23) sits within ZCode's three runs (24, 21, and 23). ZCode is cheaper per pass, $0.16 against Mouse's $0.29. On the full GLM-5.3 model, ZCode and Claude Code pass 24 to 25 tasks and cost $2 to $2.60 per pass, so Mouse on Flash comes within one or two tasks of them at a seventh to a ninth of the cost per pass.
Against Kimi K3
The same harness scored 25 of 30 on Kimi K3 at $2.79 per pass. Flash lost two tasks that K3 passed, both DeepSWE: expr-try-catch-errors and fastapi-deprecation-response-headers. It did not pass anything K3 failed. The other five misses were shared: meriyah on DeepSWE, and dna-insert, polyglot-c-py, extract-elf, and gcode-to-text on Terminal-Bench.
The cost difference is mostly the model's price. Flash lists at $0.15 per million input tokens, $0.03 cached, and $0.50 output. Cache hits covered 96.7% of input tokens, so the median passing task cost a little over a cent. The DeepSWE tasks were most of the bill: the nine of them cost $6.10 of the $6.72.
Terminal-Bench
| Task | GLM-5.3-Flash | Steps | Cost | Time | Kimi K3 |
|---|---|---|---|---|---|
| regex-log | Pass | 26 | $0.05 | 8:56 | Pass |
| openssl-selfsigned-cert | Pass | 26 | $0.01 | 3:28 | Pass |
| polyglot-c-py | Fail | 17 | $0.04 | 9:42 | Fail |
| sqlite-db-truncate | Pass | 12 | $0.01 | 3:46 | Pass |
| git-leak-recovery | Pass | 15 | $0.01 | 3:20 | Pass |
| log-summary-date-ranges | Pass | 18 | $0.01 | 3:12 | Pass |
| constraints-scheduling | Pass | 13 | $0.01 | 3:56 | Pass |
| gcode-to-text | Fail | 57 | $0.13 | 16:26 | Fail |
| dna-insert | Fail | 16 | $0.02 | 6:48 | Fail |
| largest-eigenval | Pass | 15 | $0.04 | 16:48 | Pass |
| merge-diff-arc-agi-task | Pass | 23 | $0.02 | 5:26 | Pass |
| vulnerable-secret | Pass | 16 | $0.01 | 3:16 | Pass |
| extract-elf | Fail | 21 | $0.03 | 7:02 | Fail |
| build-cython-ext | Pass | 98 | $0.17 | 12:16 | Pass |
| kv-store-grpc | Pass | 21 | $0.01 | 3:20 | Pass |
| chess-best-move | Pass | 17 | $0.01 | 4:11 | Pass |
| db-wal-recovery | Pass | 14 | $0.01 | 4:27 | Pass |
| code-from-image | Pass | 8 | $0.00 | 2:13 | Pass |
| modernize-scientific-stack | Pass | 12 | $0.01 | 2:43 | Pass |
| multi-source-data-merger | Pass | 12 | $0.01 | 3:44 | Pass |
| sanitize-git-repo | Pass | 15 | $0.01 | 3:03 | Pass |
DeepSWE
| Task | GLM-5.3-Flash | Steps | Cost | Time | Kimi K3 |
|---|---|---|---|---|---|
| anko-typed-variable-bindings | Pass | 114 | $0.38 | 20:55 | Pass |
| arktype-json-schema-refs-dependencies | Pass | 211 | $0.82 | 48:27 | Pass |
| fastapi-deprecation-response-headers | Fail | 161 | $0.54 | 42:52 | Pass |
| httpx-multipart-response-parsing | Pass | 96 | $0.34 | 31:36 | Pass |
| expr-try-catch-errors | Fail | 163 | $0.76 | 42:22 | Pass |
| python-statemachine-state-data-scoping | Pass | 250 | $1.55 | 77:48 | Pass |
| katex-multicolumn-array-spans | Pass | 127 | $0.53 | 33:43 | Pass |
| scc-bounded-memory-spilling | Pass | 100 | $0.48 | 29:46 | Pass |
| meriyah-explicit-resource-declarations | Fail | 154 | $0.71 | 39:01 | Fail |
How the run was made
- FrontierHarness's own trial driver, unmodified, at their eval commit
e837a70, on Runta with a fresh restore for every task and 4 vCPU / 8 GiB per task. Terminal-Bench through Harbor 0.22.0, DeepSWE through Pier 0.3.1. - Harness commit
315e2b8on OpenCode 1.18.27, the same checkpoint as the Kimi K3 run, with only the model changed:glm-5p3-flashserved by Fireworks. - Cost is each trial's observed tokens at Z.ai's GLM-5.3-Flash list price, which Fireworks matches. Three attempts on two tasks hit Ubuntu mirror errors while setting up the container, before the model ran; they were re-run and are not counted as failures.
- One run, 2026-09-14. The other rows were run locally by a third party, on a different provider, runtime, and harness versions (Claude Code 2.1.237, OpenCode 1.18.19, Codex 0.148.0, Pi 0.84.2), so this is a side-by-side, not a controlled ranking.
GLM results for ZCode, Claude Code, Codex, Pi, and OpenCode are from a community comparison shared this week, building on work by LotusDecoder. Mouse is built on OpenCode and is not affiliated with it or with Z.ai.