MidMarketBench

European B2B software investment benchmark

Investment judgement,benchmarked.

Not “can the model answer?” but “can the model help decide?”

Observed 18 July 20269 requested8 ranked

Methodology v0.4-mini

Observed standings from a synthetic diligence case.

8 ranked / 9 requested / 2 samples

#1Kimi K3Moonshot AI96.3Range 95.6 to 96.9Provisional tieComplete
Dimension scoresView model record
Grounding85.4
Commercial judgement98.7
Scepticism97.7
Numerical sanity100
Risk discovery96.6
Question generation96.1
European context100
Output usefulness99.7
Closed-bookn=2/2Route caveat
#2Grok 4.5xAI96.3Range 95.5 to 97.0Provisional tieComplete
Dimension scoresView model record
Grounding89.4
Commercial judgement99.1
Scepticism100
Numerical sanity100
Risk discovery99.6
Question generation96.1
European context70
Output usefulness99.7
Closed-bookn=2/2
#3GPT-5.6 SolOpenAI95.5Range 95.1 to 95.9Provisional tieComplete
Dimension scoresView model record
Grounding87.3
Commercial judgement96.9
Scepticism96.5
Numerical sanity100
Risk discovery94.2
Question generation94.5
European context100
Output usefulness99.6
Closed-bookn=2/2
#4Gemini 3.5 FlashGoogle89.1Range 86.9 to 91.3Provisional tieComplete
Dimension scoresView model record
Grounding86.7
Commercial judgement81.1
Scepticism88.8
Numerical sanity100
Risk discovery91.7
Question generation92.1
European context96.7
Output usefulness74.6
Closed-bookn=2/2
#5Claude Opus 4.8Anthropic88.8Range 80.2 to 97.4Provisional tieComplete
Dimension scoresView model record
Grounding91.3
Commercial judgement77.7
Scepticism96.5
Numerical sanity100
Risk discovery94.2
Question generation66.5
European context100
Output usefulness85.7
Closed-bookn=2/2
#6DeepSeek V4 ProDeepSeek88.6Range 87.9 to 89.4Provisional tieComplete
Dimension scoresView model record
Grounding82
Commercial judgement90.7
Scepticism87.2
Numerical sanity100
Risk discovery81.1
Question generation91.3
European context85
Output usefulness91.5
Closed-bookn=2/2
#7GLM 5.2Z.ai88.2Range 86.6 to 89.7Provisional tieComplete
Dimension scoresView model record
Grounding85.9
Commercial judgement95.9
Scepticism88.4
Numerical sanity90.7
Risk discovery83.2
Question generation87.3
European context70
Output usefulness90.4
Closed-bookn=2/2
#8Qwen3.7 MaxAlibaba65.5Range 44.6 to 86.4Complete
Dimension scoresView model record
Grounding66.5
Commercial judgement78.4
Scepticism60.8
Numerical sanity57.5
Risk discovery45.9
Question generation80.4
European context70
Output usefulness73.5
Closed-bookn=2/2
ModelRangeSamplesMode
1Kimi K3Moonshot AIProvisional tieRoute caveat96.395.6 to 96.985.498.797.7100.096.696.1100.099.72/2Closed-book
2Grok 4.5xAIProvisional tie96.395.5 to 97.089.499.1100.0100.099.696.170.099.72/2Closed-book
3GPT-5.6 SolOpenAIProvisional tie95.595.1 to 95.987.396.996.5100.094.294.5100.099.62/2Closed-book
4Gemini 3.5 FlashGoogleProvisional tie89.186.9 to 91.386.781.188.8100.091.792.196.774.62/2Closed-book
5Claude Opus 4.8AnthropicProvisional tie88.880.2 to 97.491.377.796.5100.094.266.5100.085.72/2Closed-book
6DeepSeek V4 ProDeepSeekProvisional tie88.687.9 to 89.482.090.787.2100.081.191.385.091.52/2Closed-book
7GLM 5.2Z.aiProvisional tie88.286.6 to 89.785.995.988.490.783.287.370.090.42/2Closed-book
8Qwen3.7 MaxAlibaba65.544.6 to 86.466.578.460.857.545.980.470.073.52/2Closed-book

1 requested model was excluded from ranking because a complete two-sample result was not available. Inspect the full roster

Scoring shape

Judgement carries the most weight.

A grounded but commercially naive model should not lead an investment benchmark. Overall is derived from eight dimension scores, never stored directly.

Grounding

15%

Claims are supported by packet evidence or clearly marked as inference.

Commercial judgement

20%

Identifies what matters economically and strategically to the deal.

Scepticism

15%

Challenges weak claims, inflated metrics and convenient narratives.

Numerical sanity

15%

Interprets and reconciles financial and operating data correctly.

Risk discovery

15%

Finds hidden and second-order risks across the packet.

Question generation

10%

Asks the next questions most likely to change the recommendation.

European context

5%

Accounts for fragmented markets, regulation, language and procurement.

Output usefulness

5%

Produces material a deal team can reuse with limited editing.

Interpretation

Workflow usefulness, not general intelligence.

This board reports observed OpenRouter outputs and evaluator scores for one fresh, synthetic company. It measures decision usefulness on this workflow, not general model intelligence.