Ox Alpha Benchmarks: What We Actually Know

There are no official benchmarks for Ox Alpha. The model is anonymous, there is no model card, and no lab has published scores. What exists is a small amount of third-party testing plus the capabilities stated on the OpenRouter listing. This page collects that honestly — with the caveats it deserves — and will be updated as more results appear.

What the listing claims

FocusReasoning; coding; sustained agentic work; production workloads
Context1,048,576 tokens
Max output131,072 tokens
InputText, image, video
FeaturesTool / function calling, structured JSON output, streaming

These are claims from the listing, not measurements.

Third-party results so far

Agentic coding: ~80% on a 10-task DeepSWE subset

Developer Ben Davis reported running Ox Alpha on a 10-task subset of the DeepSWE software-engineering benchmark and observed roughly an 80% pass rate. This is the most concrete number published so far, and it needs heavy qualification: ten tasks is a very small sample and the run has not been reproduced independently. Treat it as a promising anecdote, not a score.

Community reports: mixed

Early hands-on reports on Hacker News and elsewhere are mixed. Some users describe strong performance on long-context coding tasks and agentic loops; others report inconsistent or unremarkable results on general tasks. None of these reports use a controlled methodology.

How it compares to ChatGPT, Claude, Gemini, GLM and Grok

We are not going to publish a comparison table, because there are no comparable numbers to put in it. Until someone runs Ox Alpha on a standard benchmark under documented conditions, any “Ox Alpha beats X” headline is unsupported. What can be said: its 1M-token context and 131K-token output are at the upper end of what frontier model families currently offer, and its free preview pricing is obviously unmatched — for as long as it lasts.

Run your own evaluation

Because the model is free right now, the cheapest way to know whether it fits your work is to test it on your work:

  1. Pick 10–20 real tasks from your backlog (bug fixes, refactors, document analysis).
  2. Run them through the OpenRouter API with stealth/ox-alpha, and through whichever model you use today.
  3. Score outputs blind — hide which model produced which answer.
  4. Pay attention to long-context tasks specifically; that is where the 1M window should show up.

If you publish results, we would like to list them here. Use the contact page.

Try it now

Chat with Ox Alpha free in your browser, or read what Ox Alpha is and who might be behind it.

Unofficial, independent site. Not affiliated with OpenRouter or any AI lab. All trademarks belong to their owners.