Research · 2026-09-04
Model Routing Benchmarks
500 common prompt pairs · 25 canonical models · cross-fitted historical answer scores.
Score–Cost Frontier
Equivalent quality, lower cost
Laminarity’s lowest-cost operating point that met each OpenRouter router’s observed score used the same 500-prompt panel.
Against Auto
Laminarity α=0.1572
Auto: 53.2% at $0.473 / 1k. 95% score intervals: 49.0%–57.6% vs 48.8%–57.4%.
Against Auto Beta
Laminarity α=0.1336
Auto Beta: 51.8% at $0.301 / 1k. 95% score intervals: 47.6%–56.4% vs 47.6%–56.3%.
These matched points are descriptive test-set operating points selected after evaluation; the deployable headline remains Laminarity α=0.8 above. Cost deltas are Laminarity minus the comparator.
Model Benchmarks
| Fixed model | Objective score | Batch cost / 1k |
|---|---|---|
GPT-4.1 NanoOpenAI | 40.4% | $0.0628 |
GPT-5.6 LunaOpenAI | 47.7% | $0.0882 |
Gemini 2.5 Flash LiteGoogle | 48.5% | $0.203 |
Gemini 3.1 Flash LiteGoogle | 52.7% | $0.204 |
GPT-5.4 NanoOpenAI | 43.5% | $0.212 |
GPT-5.4 MiniOpenAI | 48.7% | $0.389 |
GPT-4.1 MiniOpenAI | 49.2% | $0.400 |
| 48.7% | $0.664 | |
Gemini 3.5 Flash LiteGoogle | 49.7% | $0.718 |
| 53.3% | $0.841 | |
Gemini 2.5 FlashGoogle | 55.9% | $0.936 |
Claude Haiku 4.5Anthropic | 49.0% | $0.971 |
GPT-5.6 TerraOpenAI | 53.2% | $1.09 |
Gemini 3.7 FlashGoogle | 63.2% | $1.18 |
GPT-4.1OpenAI | 51.1% | $1.18 |
| 53.1% | $1.20 | |
Claude Sonnet 5Anthropic | 53.1% | $1.98 |
Gemini 3.5 FlashGoogle | 57.2% | $2.32 |
GPT-5.6 SolOpenAI | 56.3% | $2.55 |
GPT-5.5OpenAI | 56.9% | $2.88 |
| 61.6% | $2.99 | |
Gemini 2.5 ProGoogle | 46.0% | $3.18 |
| 58.1% | $3.61 | |
Claude Opus 4.8Anthropic | 60.3% | $11 |
Claude Opus 5Anthropic | 57.1% | $11 |
The fixed rows are the 25 canonical models in the controlled OpenRouter pool. Exact-match accuracy is not reported for this replay; the score is the frozen historical primary metric.
Task Benchmarks
| Benchmark | Pairs | Laminarity | Auto | Auto Beta |
|---|---|---|---|---|
| AIME | 24 | 91.7% | 87.5% | 29.2% |
| Arena-Hard v2 | 15 | 78.3% | 81.7% | 80.0% |
| BBEH | 37 | 5.4% | 13.5% | 2.7% |
| EvalPlus | 67 | 88.1% | 80.6% | 82.1% |
| FinQA | 29 | 0.0% | 0.0% | 0.0% |
| HLCE | 49 | 40.8% | 2.0% | 2.0% |
| HLE Text MC | 27 | 55.6% | 22.2% | 7.4% |
| IFEval | 56 | 98.2% | 89.3% | 91.1% |
| IFEval Hard | 34 | 91.2% | 79.4% | 79.4% |
| MATH 500 | 43 | 93.0% | 93.0% | 88.4% |
| MMLU-Pro | 56 | 76.8% | 69.6% | 67.9% |
| OlympiadBench | 37 | 0.0% | 0.0% | 45.9% |
| SimpleQA Verified | 13 | 69.2% | 23.1% | 23.1% |
| SWE-bench Verified | 0 | — | — | — |
| WildBench | 13 | 70.0% | 60.0% | 55.4% |
Router Selection
Laminarity α=0.8
500 common prompt pairs
Gemini 3.7 Flash45190.2%
Gemini 2.5 Flash255.0%
Gemini 3.5 Flash91.8%
Claude Opus 4.881.6%
GPT-5.6 Luna40.8%
Gemini 2.5 Flash Lite20.4%
GPT-5.6 Sol10.2%
OpenRouter Auto
500 common prompt pairs
GPT-5.6 Luna24448.8%
Gemini 3.7 Flash15731.4%DeepSeek V4 Pro5911.8%
GLM 5.2285.6%
Gemini 2.5 Flash Lite102.0%
GPT-4.120.4%
OpenRouter Auto Beta
500 common prompt pairs
GPT-5.6 Luna24148.2%DeepSeek V4 Pro16933.8%
Gemini 2.5 Flash7615.2%
Gemini 2.5 Flash Lite142.8%
Default operating point
| System | Score | Cost / 1k |
|---|---|---|
| Laminarity | 63.4% | $1.15 |
| OpenRouter Auto | 53.2% | $0.473 |
| OpenRouter Auto Beta | 51.8% | $0.301 |
OpenRouter results use the completed 500-prompt selector replay. The matched-quality section above is the appropriate cost comparison.
Methodology
Laminarity was cross-fitted on the same 25-model controlled pool and evaluated on the identical stratified 500-prompt pool used for the OpenRouter Auto and Auto Beta selector replay. OpenRouter response text was not scored; each returned model ID indexed the frozen historical answer score and realized cost for that model on that prompt. Probe acquisition cost is separate from the historical cost shown here. Auto and Auto Beta are live, time-varying routers, so results apply to this frozen evaluation window.