MidMarketBench
Rank 8completeObserved OpenRouter run

Qwen3.7 Max

Model and run

Provider
Alibaba
Version
qwen3.7-max
Release
Not published
Context
1,000,000 tokens
Weights
Not published
Run status
complete
Candidate samples
2/2
Scored samples
2/2
Overall
65.5
Observed range
44.6 to 86.4
Provider-reported attributed cost
$0.411
Candidate attempts
$0.090
Judge attempts
$0.322
Median candidate latency
1.3 minutes

Model source: OpenRouter catalogue

Attributed costs include every OpenRouter-reported candidate and judge attempt for this model, including retries. The run total is authoritative for settled key spend.

OpenRouter provenance

Requested model and returned route.

Requested model
qwen/qwen3.7-max
Routed provider
Alibaba
Endpoint tag
alibaba
Returned model
qwen/qwen3.7-max
Quantisation
Not reported
Pinned route price
alibaba: $1.475 input / $4.425 output per 1M tokens

Run 2026-07-18-final / 2026-07-18 / Closed-book

Dimension shape

Where the observed score comes from.

Overall combines deterministic task checks with blinded, calibrated, cross-family judgements of the IC note.

Grounding66.5
Commercial judgement78.4
Scepticism60.8
Numerical sanity57.5
Risk discovery45.9
Question generation80.4
European context70.0
Output usefulness73.5

Complete field

Position among ranked two-sample runs.

6498

Interpretation limit

One case, 2 scored samples.

Directional mini benchmark: one fresh synthetic case and two samples per model. It is not a universal ranking of model intelligence.