There are no official benchmarks for Ox Alpha. The model is anonymous, there is no model card, and no lab has published scores. What exists is a small amount of third-party testing plus the capabilities stated on the OpenRouter listing. This page collects that honestly — with the caveats it deserves — and will be updated as more results appear.
What the listing claims
| Focus | Reasoning; coding; sustained agentic work; production workloads |
| Context | 1,048,576 tokens |
| Max output | 131,072 tokens |
| Input | Text, image, video |
| Features | Tool / function calling, structured JSON output, streaming |
These are claims from the listing, not measurements.
Third-party results so far
Agentic coding: ~80% on a 10-task DeepSWE subset
Developer Ben Davis reported running Ox Alpha on a 10-task subset of the DeepSWE software-engineering benchmark and observed roughly an 80% pass rate. This is the most concrete number published so far, and it needs heavy qualification: ten tasks is a very small sample and the run has not been reproduced independently. Treat it as a promising anecdote, not a score.
Community reports: mixed
Early hands-on reports on Hacker News and elsewhere are mixed. Some users describe strong performance on long-context coding tasks and agentic loops; others report inconsistent or unremarkable results on general tasks. None of these reports use a controlled methodology.
How it compares to ChatGPT, Claude, Gemini, GLM and Grok
We are not going to publish a comparison table, because there are no comparable numbers to put in it. Until someone runs Ox Alpha on a standard benchmark under documented conditions, any “Ox Alpha beats X” headline is unsupported. What can be said: its 1M-token context and 131K-token output are at the upper end of what frontier model families currently offer, and its free preview pricing is obviously unmatched — for as long as it lasts.
Run your own evaluation
Because the model is free right now, the cheapest way to know whether it fits your work is to test it on your work:
- Pick 10–20 real tasks from your backlog (bug fixes, refactors, document analysis).
- Run them through the OpenRouter API with
stealth/ox-alpha, and through whichever model you use today. - Score outputs blind — hide which model produced which answer.
- Pay attention to long-context tasks specifically; that is where the 1M window should show up.
If you publish results, we would like to list them here. Use the contact page.
Try it now
Chat with Ox Alpha free in your browser, or read what Ox Alpha is and who might be behind it.