Mouse on FrontierHarness
Mouse scored 24/30 with the highest pass rate.
FrontierHarness is a benchmark that runs the same model, Kimi K3, through different coding agents on the same 30 tasks. Mouse passed 24 of 30, the highest pass rate and the fastest median time of any harness on the board.
The benchmark
The 30 tasks come in two sets. Twenty-one are from Terminal-Bench, short jobs done at a command line, like recovering a corrupted database or building a software package. The other nine are DeepSWE tasks: real bug fixes and features in real open source projects, graded by tests the agent never sees.
The model is fixed. What changes is the harness, the program that gives the model its instructions, its tools, and the judgment of when a task is done.
How Mouse works
The Mouse agent harness is built on top of the OpenCode engine. We designed our system prompt, configuration, rules, and set of skills to steer a custom loop that outperforms on coding tasks.
The loop is the biggest differentiator within Mouse. Most agents finish when the model decides a task is done. With Mouse, we use a deterministic approach that mandates the workspace inspects, verifies, and ensures the evidence of the completed task holds up. This loop iterates until the desired outcome.
This approach shows up when long-horizon jobs are required, which is why OpenCode passed 0 of the DeepSWE tasks and Mouse passed 6.
This is the same loop that powers Night Shift, how we run agents on long time horizon jobs and overnight.
Leaderboard
| Harness | Pass | Terminal | DeepSWE | $ / pass | Median $ | Cache | Time |
|---|---|---|---|---|---|---|---|
| 24/30 | 18/21 | 6/9 | $3.13 | $0.31 | 84% | 4:10 | |
| Codex | 20/30 | 15/21 | 5/9 | $3.47 | $0.12 | 88% | 6:43 |
| DSH Creator | 19/30 | 15/21 | 4/9 | $3.28 | $0.12 | 84% | 6:43 |
| Claude Code | 19/30 | 16/21 | 3/9 | $18.34 | $0.29 | 68% | 9:37 |
| Pi | 18/30 | 16/21 | 2/9 | $2.43 | $0.07 | 79% | 7:33 |
| DSH Standard | 18/30 | 16/21 | 2/9 | $3.46 | $0.12 | 86% | 6:16 |
| DSH PTC | 18/30 | 16/21 | 2/9 | $4.58 | $0.14 | 87% | 7:43 |
| 17/30 | 14/21 | 3/9 | $3.65 | $0.18 | 88% | 7:55 | |
| DSH Minimal | 17/30 | 15/21 | 2/9 | $4.72 | $0.12 | 85% | 5:41 |
| Oh My Pi | 17/30 | 16/21 | 1/9 | $4.75 | $0.14 | 82% | 6:46 |
| 16/30 | 16/21 | 0/9 | $1.05 | $0.07 | 70% | 6:17 | |
| Hermes | 15/30 | 13/21 | 2/9 | $2.90 | $0.17 | 86% | 6:57 |
| OpenCode | 15/30 | 15/21 | 0/9 | $3.24 | $0.06 | 78% | 6:27 |
Cost per pass is total spend on all 30 tasks divided by passes. Median cost, cache rate, and time cover passing tasks only. Mouse was the only harness to solve kv-store-grpc and scc-bounded-memory-spilling.
Cost
Mouse is fourth on cost per pass and last on median cost. Each verification round is another model call, which increases the harness's costs. Our next move is to improve cache and compaction to bring these costs down.
Terminal-Bench
| Task | Mouse | Steps | Cost | Time | Field |
|---|---|---|---|---|---|
| regex-log | Pass | 21 | $0.51 | 7:50 | 11/12 |
| openssl-selfsigned-cert | Pass | 13 | $0.20 | 2:44 | 12/12 |
| polyglot-c-py | Pass | 14 | $0.15 | 2:35 | 12/12 |
| sqlite-db-truncate | Pass | 21 | $0.32 | 4:17 | 12/12 |
| git-leak-recovery | Pass | 15 | $0.14 | 2:06 | 12/12 |
| log-summary-date-ranges | Pass | 21 | $0.25 | 3:32 | 12/12 |
| constraints-scheduling | Pass | 15 | $0.31 | 4:49 | 12/12 |
| gcode-to-text | Fail | 27 | $0.71 | 9:54 | 5/12 |
| dna-insert | Pass | 18 | $0.71 | 11:05 | 2/12 |
| largest-eigenval | Fail | 20 | $1.03 | – | 0/12 |
| merge-diff-arc-agi-task | Pass | 19 | $0.30 | 4:03 | 12/12 |
| vulnerable-secret | Pass | 17 | $0.18 | 2:15 | 12/12 |
| extract-elf | Fail | 13 | $0.49 | 7:49 | 2/12 |
| build-cython-ext | Pass | 47 | $0.93 | 13:08 | 12/12 |
| kv-store-grpc | Pass | 17 | $0.15 | 2:32 | 0/12 |
| chess-best-move | Pass | 18 | $0.59 | 6:13 | 5/12 |
| db-wal-recovery | Pass | 13 | $0.15 | 2:23 | 12/12 |
| code-from-image | Pass | 8 | $0.09 | 1:09 | 4/12 |
| modernize-scientific-stack | Pass | 10 | $0.21 | 3:27 | 12/12 |
| multi-source-data-merger | Pass | 10 | $0.19 | 3:07 | 12/12 |
| sanitize-git-repo | Pass | 26 | $0.47 | 3:58 | 10/12 |
Field is how many of the 12 published harnesses passed the task.
DeepSWE
| Task | Mouse | Steps | Cost | Time | Field |
|---|---|---|---|---|---|
| anko-typed-variable-bindings | Fail | 52 | $2.64 | 21:01 | 4/12 |
| arktype-json-schema-refs-dependencies | Pass | 190 | $10.32 | 89:36 | 2/12 |
| fastapi-deprecation-response-headers | Pass | 150 | $7.08 | 85:32 | 4/12 |
| httpx-multipart-response-parsing | Pass | 141 | $6.06 | 70:48 | 5/12 |
| expr-try-catch-errors | Pass | 134 | $8.59 | 44:43 | 2/12 |
| python-statemachine-state-data-scoping | Pass | 201 | $13.49 | 73:21 | 7/12 |
| katex-multicolumn-array-spans | Fail | 117 | $5.65 | 44:52 | 1/12 |
| scc-bounded-memory-spilling | Pass | 84 | $4.65 | 52:44 | 0/12 |
| meriyah-explicit-resource-declarations | Fail | 139 | $8.47 | 65:49 | 1/12 |
Two of the three misses passed nearly all of their hidden tests, 92 of 94 and 46 of 49.
Setup
- Harbor 0.22, the benchmark's runner, on a Google Cloud VM with 4 tasks in parallel.
- Kimi K3 through OpenRouter, pinned to the Fireworks servers the benchmark uses.
- One trial per task. Cost at list price. Time is wall clock per task.
Mouse is built on top of OpenCode. Run date: September 3, 2026.

