Model index
Observed model behaviour through one investment lens.
8 models completed 2 scored OpenRouter samples against the same synthetic case. Incomplete or unavailable requests are disclosed below and excluded from ranking.
Complete runs
Ranked on complete evidence.
| Rank | Model | Provider | Overall | Range | Samples | OpenRouter route | Reported attributed cost | Median latency |
|---|---|---|---|---|---|---|---|---|
| 1 | MKimi K3Route caveat | Moonshot AI | 96.3 | 95.6 to 96.9 | 2/2 | Moonshot AI / INT4 | $0.894 | 6.4m |
| 2 | Grok 4.5 | xAI | 96.3 | 95.5 to 97.0 | 2/2 | xAI | $0.328 | 1.5m |
| 3 | OAIGPT-5.6 Sol | OpenAI | 95.5 | 95.1 to 95.9 | 2/2 | OpenAI | $0.559 | 1.7m |
| 4 | Gemini 3.5 Flash | 89.1 | 86.9 to 91.3 | 2/2 | $0.725 | 2.0m | ||
| 5 | Claude Opus 4.8 | Anthropic | 88.8 | 80.2 to 97.4 | 2/2 | Azure | $0.649 | 1.3m |
| 6 | DeepSeek V4 Pro | DeepSeek | 88.6 | 87.9 to 89.4 | 2/2 | StreamLake / FP8 | $0.361 | 3.1m |
| 7 | ZGLM 5.2 | Z.ai | 88.2 | 86.6 to 89.7 | 2/2 | StreamLake / FP8 | $0.261 | 2.3m |
| 8 | Qwen3.7 Max | Alibaba | 65.5 | 44.6 to 86.4 | 2/2 | Alibaba | $0.411 | 1.3m |
Not ranked
Unavailable or incomplete runs.
Infrastructure and routing failures are not converted into zero scores. These models remain in the run ledger without a rank.
| Model | Provider | Status | Candidate samples | OpenRouter route | Recorded reason |
|---|---|---|---|---|---|
| Claude Fable 5 | Anthropic | unavailable | 0/2 | anthropic/claude-fable-5 | Claude Fable 5 is not available. Learn more: https://www.anthropic.com/news/fable-mythos-access |