Measured, not claimed.
Every number below came out of our own coding agent driving an open-weight model, scored with the standard SWE-bench evaluator. We publish the runs we lose as well as the ones we win.
On Lite, we lead.
GLM 5.3 Flash on our harness resolves 63.0% (189/300) — ahead of MiniMax M2.5’s native 56.3% and of Claude Opus 4.6, a closed frontier model. The same harness with MiniMax M2.5 resolves 51.3%, which is the honest spread: the harness is model-agnostic, and the model you put in it still matters.
On Verified, we trail — and we say so.
Our harness drove MiniMax M2.5 to 70.0% (350/500). The same model reports 80.2% on its own scaffolding, so roughly ten points here are ours to close, not the model’s. That gap is scaffold work — reproduce-test-first, regression filtering, best-of-N, hierarchical localization — and it is the work we are doing.
The harness is the variable.
Two frontier models, one harness
GLM 5.3 Flash and MiniMax M2.5 both run on the same coding agent, with no re-integration. Swapping the model is a configuration change, not a rebuild.
No foreign kill switch
Every model we run is open weight, so no vendor policy change can take it away. The closed frontier scores higher today; it is also the thing that can go dark.
Raw counts, standard scorer
We publish resolved/total, not just a percentage, and score with the standard SWE-bench evaluator. Numbers you can re-run are the only kind worth quoting.
SWE-bench Lite and Verified are different datasets and are never merged into one chart. Reference scores for other models are their own published or third-party leaderboard results, shown so the harness-versus-model gap is visible rather than hidden.
