Benchmarks
Benchmarks & methodology.
How Ox Alpha lines up against the current frontier — the full table, the caveats behind every number, and exactly how we put it together.
Last updated: August 2026
Most comparison tables give you a verdict without the receipts. This one shows the numbers and then tells you where they come from, what they can't tell you, and why the free preview changes the maths.
Ox Alpha is a stealth preview model, so treat its figures as directional rather than final — the honest way to read this page is "frontier-adjacent quality at a price no shipping model can match."
The full table
Nine dimensions, five models, one column that's free.
| Metric | Ox Alpha | Fable 5 | GLM-5 | GPT-5.6 | Grok 4.6 |
|---|---|---|---|---|---|
| Intelligence Index | 59* | 62 | 57 | 59 | 61 |
| Agentic coding | 78* | 79 | 72 | 80 | 74 |
| Context window | 1M | 1M | 200K | 400K | 500K |
| Max output | 131K | 64K | 128K | 128K | 64K |
| Throughput (tps) | 24* | 63 | 90 | 70 | 120 |
| Price / 1M out | $0 | $50 | $1.50 | $15 | $6 |
| Multimodal | |||||
| Open weights | |||||
| Free to use |
Composite score across reasoning, knowledge and math benchmarks (higher is better).
Indicative agentic-coding score — real-world, multi-file tasks with tools (higher is better).
Maximum tokens the model can attend to in a single request.
Largest response the model will generate in one call.
Median output speed in tokens per second (higher is faster).
List price per million output tokens (lower is cheaper).
Accepts image input alongside text.
Weights can be downloaded and self-hosted.
Usable at no cost right now.
Head-to-head
Want a focused, one-on-one breakdown? Pick a match-up.
Methodology
Intelligence Index is a composite of public reasoning, knowledge and math benchmarks in the spirit of the Artificial Analysis Intelligence Index. Agentic coding reflects real-world, multi-file tasks executed with tools rather than isolated snippets.
Context window, max output, throughput and price are provider-reported figures. Price is the list rate per one million output tokens; input tokens are usually cheaper and are omitted to keep the table readable.
Ox Alpha is served anonymously in preview, so its figures (marked *) are community-reported estimates rather than a vendor datasheet. They will move as the model is tuned — we refresh this page as better data appears.
Numbers change fast in this space. Use this as a directional guide for picking a model, not as a benchmark leaderboard of record. When in doubt, run your own eval on your own prompts.
See it for yourself
Numbers are one thing. Hand Ox Alpha a real task and judge the output.