Ox Alpha was revealed as Z.ai GLM-5.3-Flash on August 26, 2026. Read the full story →

Ox Alpha’s 80% Coding Benchmark Has Been Corrected: ~63% on the Full DeepSWE Set

Historical note — updated August 27, 2026: Ox Alpha was revealed by Z.ai as GLM-5.3-Flash on August 26, 2026. This article preserves the information available at publication. See the current model reference and reveal report for current specifications, pricing, and availability.

The number that made Ox Alpha famous has been revised by the person who produced it. Developer Ben Davis, whose August 21 post reported the anonymous stealth model passing 8 of 10 DeepSWE tasks — “80%”, ahead of a Claude-family model at 65% and a ChatGPT-family model at 52% on the same subset — has since completed a run of the full 113-task DeepSWE set. The result, as relayed in widely shared posts citing him, is roughly 63%: “more or less on par” with a mid-tier ChatGPT-family model, not ahead of the field.

Why the two numbers differ

Ten tasks is under 9% of DeepSWE. Davis said so in the original post — “there could be a ton of variance in its real score, this is a subset” — but the caveat did not travel as far as the headline. On a 10-task sample, one task is 10 points; a two-task swing is the entire gap between “tops the table” and “mid-pack”. The full run removes that noise, which is why Davis described the ~63% figure as the one that makes more sense.

What DeepSWE measures

DeepSWE is a July 2026 benchmark of 113 original, long-horizon software-engineering tasks across 91 live open-source repositories, with hand-written verifiers and reference solutions kept out of public training data. It is deliberately harder and less contaminated than older coding benchmarks — so a mid-60s score from a free, anonymous preview is still a respectable result. It is just not the “beats everything” result that circulated.

What this changes

  • Our benchmarks page now leads with the full-run figure and keeps the 10-task result as history, with both caveats spelled out.
  • Pages elsewhere that still headline 80% — including per-task tables built from the subset — are reporting a superseded number. Check the sample size before you trust any Ox Alpha score.
  • Nothing here is official. There is still no model card, no vendor benchmark, and no leaderboard-protocol run. The identity of the developer also remains unconfirmed.

The cheapest benchmark is still your own backlog. Try it in the browser or follow the API guide. We have not seen Davis’s raw logs; if a reproducible run is published, we will link it here.

Postscript, August 26, 2026: Z.ai revealed the model as GLM-5.3-Flash and published its own DeepSWE v1.1 figure — 63.4. The correction reported here was right, and the viral 80% never was. Full benchmark page →

Sources and verification

Last substantively reviewed: August 27, 2026. Reviewed by: OxAlpha.chat Editorial Team.

How we verified this page

This correction distinguishes an early partial-sample estimate from the full benchmark result. We retain both numbers with their sample sizes and label Z.ai scores as vendor-reported.

Primary sources

Found a factual error or a changed price? Send us the page URL and supporting source; corrections follow our editorial policy.

Unofficial, independent site. Not affiliated with OpenRouter or any AI lab. All trademarks belong to their owners.