Research · 2026-09-04

Model Routing Benchmarks

500 common prompt pairs · 25 canonical models · cross-fitted historical answer scores.

500common prompt pairs
25canonical models
63.4%default policy score
$1.15batch cost / 1k

Score–Cost Frontier

Higher and farther left is better. The dashed line connects fixed-model Pareto points; the purple line shows Laminarity operating points. Costs use a logarithmic scale.

Equivalent quality, lower cost

Laminarity’s lowest-cost operating point that met each OpenRouter router’s observed score used the same 500-prompt panel.

Against Auto

Laminarity α=0.1572

16.2% less
53.3%Laminarity score
$0.396Laminarity cost / 1k
+0.05 ppscore delta vs Auto
$0.0769cost saved / 1k

Auto: 53.2% at $0.473 / 1k. 95% score intervals: 49.0%–57.6% vs 48.8%–57.4%.

Against Auto Beta

Laminarity α=0.1336

20.7% less
51.9%Laminarity score
$0.239Laminarity cost / 1k
+0.08 ppscore delta vs Auto Beta
$0.0623cost saved / 1k

Auto Beta: 51.8% at $0.301 / 1k. 95% score intervals: 47.6%–56.4% vs 47.6%–56.3%.

These matched points are descriptive test-set operating points selected after evaluation; the deployable headline remains Laminarity α=0.8 above. Cost deltas are Laminarity minus the comparator.

Model Benchmarks

Fixed modelObjective scoreBatch cost / 1k
GPT-4.1 NanoOpenAI40.4%$0.0628
GPT-5.6 LunaOpenAI47.7%$0.0882
Gemini 2.5 Flash LiteGoogle48.5%$0.203
Gemini 3.1 Flash LiteGoogle52.7%$0.204
GPT-5.4 NanoOpenAI43.5%$0.212
GPT-5.4 MiniOpenAI48.7%$0.389
GPT-4.1 MiniOpenAI49.2%$0.400
GLM 5.2DeepInfra48.7%$0.664
Gemini 3.5 Flash LiteGoogle49.7%$0.718
DeepSeek V4 Pro 0813Fireworks AI53.3%$0.841
Gemini 2.5 FlashGoogle55.9%$0.936
Claude Haiku 4.5Anthropic49.0%$0.971
GPT-5.6 TerraOpenAI53.2%$1.09
Gemini 3.7 FlashGoogle63.2%$1.18
GPT-4.1OpenAI51.1%$1.18
DeepSeek V4 ProDeepInfra53.1%$1.20
Claude Sonnet 5Anthropic53.1%$1.98
Gemini 3.5 FlashGoogle57.2%$2.32
GPT-5.6 SolOpenAI56.3%$2.55
GPT-5.5OpenAI56.9%$2.88
Muse Spark 1.2Meta61.6%$2.99
Gemini 2.5 ProGoogle46.0%$3.18
Kimi K3Fireworks AI58.1%$3.61
Claude Opus 4.8Anthropic60.3%$11
Claude Opus 5Anthropic57.1%$11

The fixed rows are the 25 canonical models in the controlled OpenRouter pool. Exact-match accuracy is not reported for this replay; the score is the frozen historical primary metric.

Task Benchmarks

BenchmarkPairsLaminarityAutoAuto Beta
AIME2491.7%87.5%29.2%
Arena-Hard v21578.3%81.7%80.0%
BBEH375.4%13.5%2.7%
EvalPlus6788.1%80.6%82.1%
FinQA290.0%0.0%0.0%
HLCE4940.8%2.0%2.0%
HLE Text MC2755.6%22.2%7.4%
IFEval5698.2%89.3%91.1%
IFEval Hard3491.2%79.4%79.4%
MATH 5004393.0%93.0%88.4%
MMLU-Pro5676.8%69.6%67.9%
OlympiadBench370.0%0.0%45.9%
SimpleQA Verified1369.2%23.1%23.1%
SWE-bench Verified0
WildBench1370.0%60.0%55.4%

Router Selection

Laminarity α=0.8

500 common prompt pairs

7 models
  • Gemini 3.7 Flash45190.2%
  • Gemini 2.5 Flash255.0%
  • Gemini 3.5 Flash91.8%
  • Claude Opus 4.881.6%
  • GPT-5.6 Luna40.8%
  • Gemini 2.5 Flash Lite20.4%
  • GPT-5.6 Sol10.2%

OpenRouter Auto

500 common prompt pairs

6 models
  • GPT-5.6 Luna24448.8%
  • Gemini 3.7 Flash15731.4%
  • DeepSeek V4 Pro5911.8%
  • GLM 5.2285.6%
  • Gemini 2.5 Flash Lite102.0%
  • GPT-4.120.4%

OpenRouter Auto Beta

500 common prompt pairs

4 models
  • GPT-5.6 Luna24148.2%
  • DeepSeek V4 Pro16933.8%
  • Gemini 2.5 Flash7615.2%
  • Gemini 2.5 Flash Lite142.8%

Default operating point

SystemScoreCost / 1k
Laminarity63.4%$1.15
OpenRouter Auto53.2%$0.473
OpenRouter Auto Beta51.8%$0.301

OpenRouter results use the completed 500-prompt selector replay. The matched-quality section above is the appropriate cost comparison.

Methodology

Laminarity was cross-fitted on the same 25-model controlled pool and evaluated on the identical stratified 500-prompt pool used for the OpenRouter Auto and Auto Beta selector replay. OpenRouter response text was not scored; each returned model ID indexed the frozen historical answer score and realized cost for that model on that prompt. Probe acquisition cost is separate from the historical cost shown here. Auto and Auto Beta are live, time-varying routers, so results apply to this frozen evaluation window.

Download homepage benchmark data