Grounding
15%Claims are supported by packet evidence or clearly marked as inference.
European B2B software investment benchmark
Not “can the model answer?” but “can the model help decide?”
Methodology v0.4-mini
8 ranked / 9 requested / 2 samples
| Model | Range | Samples | Mode | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | MKimi K3Moonshot AIProvisional tieRoute caveat | 96.3 | 95.6 to 96.9 | 85.4 | 98.7 | 97.7 | 100.0 | 96.6 | 96.1 | 100.0 | 99.7 | 2/2 | Closed-book |
| 2 | Grok 4.5xAIProvisional tie | 96.3 | 95.5 to 97.0 | 89.4 | 99.1 | 100.0 | 100.0 | 99.6 | 96.1 | 70.0 | 99.7 | 2/2 | Closed-book |
| 3 | OAIGPT-5.6 SolOpenAIProvisional tie | 95.5 | 95.1 to 95.9 | 87.3 | 96.9 | 96.5 | 100.0 | 94.2 | 94.5 | 100.0 | 99.6 | 2/2 | Closed-book |
| 4 | Gemini 3.5 FlashGoogleProvisional tie | 89.1 | 86.9 to 91.3 | 86.7 | 81.1 | 88.8 | 100.0 | 91.7 | 92.1 | 96.7 | 74.6 | 2/2 | Closed-book |
| 5 | Claude Opus 4.8AnthropicProvisional tie | 88.8 | 80.2 to 97.4 | 91.3 | 77.7 | 96.5 | 100.0 | 94.2 | 66.5 | 100.0 | 85.7 | 2/2 | Closed-book |
| 6 | DeepSeek V4 ProDeepSeekProvisional tie | 88.6 | 87.9 to 89.4 | 82.0 | 90.7 | 87.2 | 100.0 | 81.1 | 91.3 | 85.0 | 91.5 | 2/2 | Closed-book |
| 7 | ZGLM 5.2Z.aiProvisional tie | 88.2 | 86.6 to 89.7 | 85.9 | 95.9 | 88.4 | 90.7 | 83.2 | 87.3 | 70.0 | 90.4 | 2/2 | Closed-book |
| 8 | Qwen3.7 MaxAlibaba | 65.5 | 44.6 to 86.4 | 66.5 | 78.4 | 60.8 | 57.5 | 45.9 | 80.4 | 70.0 | 73.5 | 2/2 | Closed-book |
1 requested model was excluded from ranking because a complete two-sample result was not available. Inspect the full roster
Scoring shape
A grounded but commercially naive model should not lead an investment benchmark. Overall is derived from eight dimension scores, never stored directly.
Claims are supported by packet evidence or clearly marked as inference.
Identifies what matters economically and strategically to the deal.
Challenges weak claims, inflated metrics and convenient narratives.
Interprets and reconciles financial and operating data correctly.
Finds hidden and second-order risks across the packet.
Asks the next questions most likely to change the recommendation.
Accounts for fragmented markets, regulation, language and procurement.
Produces material a deal team can reuse with limited editing.
Interpretation
This board reports observed OpenRouter outputs and evaluator scores for one fresh, synthetic company. It measures decision usefulness on this workflow, not general model intelligence.