XRB-1000
Xantly Routing Benchmark v1, September 2026. 1,000 questions to xantly/auto with no model pinned, against the ten most accurate models pinned on the same questions. Every number measured on the platform.
XRB-1000 is the Xantly Routing Benchmark: 1,000 questions from eight public datasets, sent once each to xantly/auto with no model pinned, and answered by every model in the fleet pinned on the same questions. Every number is measured on the Xantly platform; nothing is a published score or an estimate.
Accuracy and cost against the top models
Xantly auto answered 86.3% for $1.72. Claude Opus 4.6 reached the same accuracy for 4.6 times the cost. Gemini 3.5 Flash scored 1.1 points higher for 5.3 times the cost, and Gemini 3.1 Pro scored 0.4 points lower for 6.6 times the cost.
| Model | Accuracy | Cost, all 1,000 | Cost relative to Xantly auto |
|---|---|---|---|
| Gemini 3.5 Flash | 87.4% | $9.17 | 5.3x |
| Claude Opus 4.6 | 86.3% | $7.85 | 4.6x |
| Xantly auto | 86.3% | $1.72 | 1.0x |
| Gemini 3.1 Pro Preview | 85.9% | $11.42 | 6.6x |
| Claude Opus 4.5 | 84.7% | $8.90 | 5.2x |
| Claude Sonnet 4.6 | 82.5% | $4.11 | 2.4x |
| Kimi K2.5 | 81.1% | $1.40 | 0.8x |
| Claude Haiku 4.5 | 79.4% | $1.91 | 1.1x |
| DeepSeek V3.2 | 79.0% | $1.15 | 0.7x |
| Qwen3 Next 80B | 78.9% | $0.94 | 0.5x |
| Gemini 3.1 Flash Lite | 78.7% | $0.28 | 0.2x |
The ten most accurate pinned models and the router, on the same 1,000 questions. Cost is the provider cost for all 1,000 answers; the multiple is that cost divided by the router's.
Setup
The suite
1,000 questions drawn from eight public datasets, fixed once and reused for every model. Difficulty labels come from the suite, not from the router.
| Part | Source | Difficulty | Questions |
|---|---|---|---|
| Commonsense | HellaSwag | trivial | 100 |
| Science questions | ARC | trivial | 100 |
| Grade-school math | GSM8K | easy | 200 |
| Knowledge | MMLU-Pro | medium | 150 |
| Code | HumanEval | medium | 100 |
| Competition math | MATH-500 | hard | 200 |
| Expert questions | Humanity's Last Exam | extreme | 120 |
| Olympiad math | AIME 2024 | extreme | 30 |
| Total | 1,000 |
How each system was run
- The router. Each question was sent once to
xantly/autothrough the public API, exactly as a customer sends it: no model pinned, no routing hints, streaming on, one call per question. The gateway chose the model per request. - The pinned models. Each comparison model answered the same 1,000 questions through the same API with the model pinned. The ten most accurate are shown above.
- Grading. One automatic grader scored every answer for every system: exact match for numeric and multiple-choice answers, test execution for code, and the dataset's own answer key throughout.
- Cost. Provider list price at the tokens each model actually used, taken from the gateway's billing records. The router's cost is the sum of what its chosen models charged, including any call the gateway made on the caller's behalf.
Results by part of the suite
| Part | Difficulty | Questions | Router correct | Router accuracy |
|---|---|---|---|---|
| Commonsense (HellaSwag) | trivial | 100 | 85 | 85.0% |
| Science questions (ARC) | trivial | 100 | 99 | 99.0% |
| Grade-school math (GSM8K) | easy | 200 | 192 | 96.0% |
| Knowledge (MMLU-Pro) | medium | 150 | 132 | 88.0% |
| Code (HumanEval) | medium | 100 | 97 | 97.0% |
| Competition math (MATH-500) | hard | 200 | 181 | 90.5% |
| Expert questions (Humanity's Last Exam) | extreme | 120 | 54 | 45.0% |
| Olympiad math (AIME 2024) | extreme | 30 | 23 | 76.7% |
| Total | 1,000 | 863 | 86.3% |
Reliability
- 1,000 of 1,000 requests returned an answer. No request timed out, failed over on an error, or returned empty.
- Twice, a thinking model spent its whole output budget reasoning and produced no answer. The gateway detected it, did not serve the empty reply, and asked another model; the caller received a normal answer. Both requests are counted at their full cost.
Reproducibility
- The suite, the grader and the cost accounting are fixed; the same questions and the same grader scored every system on this page.
- The run's records (every request, the served model, the graded answer and the billed cost) are retained, so the page can be regenerated from them.
- A new run replaces this page when the platform changes. The revision history records each one.
Revision history
| Run | Router accuracy | Router cost | Note |
|---|---|---|---|
| September 2026 | 86.3% | $1.72 | First published run of XRB-1000 v1 |
Every model the router used
All 20 models the router chose across the 1,000 requests. Accuracy per row is on the questions that model received and is not comparable between rows: the router sends the hard questions to the strong models and the easy ones to the cheap models, so the strong models carry the hard misses by design.
| Model served | Requests | Share | Correct | Accuracy | Cost | Cost per request |
|---|---|---|---|---|---|---|
| Gemini 3.7 Flash | 245 | 24.5% | 192 | 78.4% | $0.703 | $0.0029 |
| GPT-OSS 20B | 185 | 18.5% | 179 | 96.8% | $0.022 | $0.0001 |
| GPT-OSS 120B | 128 | 12.8% | 123 | 96.1% | $0.026 | $0.0002 |
| Gemini 3.1 Flash Lite | 83 | 8.3% | 74 | 89.2% | $0.021 | $0.0002 |
| Kimi K2.5 | 61 | 6.1% | 45 | 73.8% | $0.153 | $0.0025 |
| Qwen3 Next 80B | 54 | 5.4% | 52 | 96.3% | $0.022 | $0.0004 |
| DeepSeek V3.2 | 52 | 5.2% | 51 | 98.1% | $0.058 | $0.0011 |
| Gemini 3.5 Flash | 44 | 4.4% | 36 | 81.8% | $0.479 | $0.0109 |
| MiniMax M2.5 | 32 | 3.2% | 23 | 71.9% | $0.037 | $0.0011 |
| Gemini 3.5 Flash Lite | 30 | 3.0% | 24 | 80.0% | $0.015 | $0.0005 |
| Gemini 3.6 Flash | 17 | 1.7% | 5 | 29.4% | $0.087 | $0.0051 |
| Gemma 3 4B | 17 | 1.7% | 14 | 82.4% | $0.000 | $0.0000 |
| Claude Sonnet 4.6 | 13 | 1.3% | 11 | 84.6% | $0.028 | $0.0021 |
| Nemotron Super 3 120B | 12 | 1.2% | 10 | 83.3% | $0.003 | $0.0002 |
| Gemma 3 12B | 10 | 1.0% | 10 | 100.0% | $0.000 | $0.0000 |
| Claude Opus 4.6 | 4 | 0.4% | 2 | 50.0% | $0.045 | $0.0113 |
| Kimi K2 Thinking | 4 | 0.4% | 4 | 100.0% | $0.024 | $0.0059 |
| Gemma 3 27B | 4 | 0.4% | 4 | 100.0% | $0.000 | $0.0001 |
| GLM-5 | 3 | 0.3% | 2 | 66.7% | $0.001 | $0.0003 |
| Llama 3.1 70B | 2 | 0.2% | 2 | 100.0% | $0.001 | $0.0003 |
| Total, xantly/auto | 1,000 | 100% | 863 | 86.3% | $1.724 | $0.0017 |