XRB-1000

Xantly Routing Benchmark v1, September 2026. 1,000 questions to xantly/auto with no model pinned, against the ten most accurate models pinned on the same questions. Every number measured on the platform.

XRB-1000 is the Xantly Routing Benchmark: 1,000 questions from eight public datasets, sent once each to xantly/auto with no model pinned, and answered by every model in the fleet pinned on the same questions. Every number is measured on the Xantly platform; nothing is a published score or an estimate.

Accuracy and cost against the top models

Xantly auto answered 86.3% for $1.72. Claude Opus 4.6 reached the same accuracy for 4.6 times the cost. Gemini 3.5 Flash scored 1.1 points higher for 5.3 times the cost, and Gemini 3.1 Pro scored 0.4 points lower for 6.6 times the cost.

ModelAccuracyCost, all 1,000Cost relative to Xantly auto
Gemini 3.5 Flash87.4%$9.175.3x
Claude Opus 4.686.3%$7.854.6x
Xantly auto86.3%$1.721.0x
Gemini 3.1 Pro Preview85.9%$11.426.6x
Claude Opus 4.584.7%$8.905.2x
Claude Sonnet 4.682.5%$4.112.4x
Kimi K2.581.1%$1.400.8x
Claude Haiku 4.579.4%$1.911.1x
DeepSeek V3.279.0%$1.150.7x
Qwen3 Next 80B78.9%$0.940.5x
Gemini 3.1 Flash Lite78.7%$0.280.2x

The ten most accurate pinned models and the router, on the same 1,000 questions. Cost is the provider cost for all 1,000 answers; the multiple is that cost divided by the router's.

Setup

The suite

1,000 questions drawn from eight public datasets, fixed once and reused for every model. Difficulty labels come from the suite, not from the router.

PartSourceDifficultyQuestions
CommonsenseHellaSwagtrivial100
Science questionsARCtrivial100
Grade-school mathGSM8Keasy200
KnowledgeMMLU-Promedium150
CodeHumanEvalmedium100
Competition mathMATH-500hard200
Expert questionsHumanity's Last Examextreme120
Olympiad mathAIME 2024extreme30
Total1,000

How each system was run

Results by part of the suite

PartDifficultyQuestionsRouter correctRouter accuracy
Commonsense (HellaSwag)trivial1008585.0%
Science questions (ARC)trivial1009999.0%
Grade-school math (GSM8K)easy20019296.0%
Knowledge (MMLU-Pro)medium15013288.0%
Code (HumanEval)medium1009797.0%
Competition math (MATH-500)hard20018190.5%
Expert questions (Humanity's Last Exam)extreme1205445.0%
Olympiad math (AIME 2024)extreme302376.7%
Total1,00086386.3%

Reliability

Reproducibility

Revision history

RunRouter accuracyRouter costNote
September 202686.3%$1.72First published run of XRB-1000 v1

Every model the router used

All 20 models the router chose across the 1,000 requests. Accuracy per row is on the questions that model received and is not comparable between rows: the router sends the hard questions to the strong models and the easy ones to the cheap models, so the strong models carry the hard misses by design.

Model servedRequestsShareCorrectAccuracyCostCost per request
Gemini 3.7 Flash24524.5%19278.4%$0.703$0.0029
GPT-OSS 20B18518.5%17996.8%$0.022$0.0001
GPT-OSS 120B12812.8%12396.1%$0.026$0.0002
Gemini 3.1 Flash Lite838.3%7489.2%$0.021$0.0002
Kimi K2.5616.1%4573.8%$0.153$0.0025
Qwen3 Next 80B545.4%5296.3%$0.022$0.0004
DeepSeek V3.2525.2%5198.1%$0.058$0.0011
Gemini 3.5 Flash444.4%3681.8%$0.479$0.0109
MiniMax M2.5323.2%2371.9%$0.037$0.0011
Gemini 3.5 Flash Lite303.0%2480.0%$0.015$0.0005
Gemini 3.6 Flash171.7%529.4%$0.087$0.0051
Gemma 3 4B171.7%1482.4%$0.000$0.0000
Claude Sonnet 4.6131.3%1184.6%$0.028$0.0021
Nemotron Super 3 120B121.2%1083.3%$0.003$0.0002
Gemma 3 12B101.0%10100.0%$0.000$0.0000
Claude Opus 4.640.4%250.0%$0.045$0.0113
Kimi K2 Thinking40.4%4100.0%$0.024$0.0059
Gemma 3 27B40.4%4100.0%$0.000$0.0001
GLM-530.3%266.7%$0.001$0.0003
Llama 3.1 70B20.2%2100.0%$0.001$0.0003
Total, xantly/auto1,000100%86386.3%$1.724$0.0017