MidMarketBench
Rank 6completeObserved OpenRouter runProvisional tie

DeepSeek V4 Pro

Model and run

Provider
DeepSeek
Version
deepseek-v4-pro
Release
Not published
Context
1,048,576 tokens
Weights
Not published
Run status
complete
Candidate samples
2/2
Scored samples
2/2
Overall
88.6
Observed range
87.9 to 89.4
Provider-reported attributed cost
$0.361
Candidate attempts
$0.039
Judge attempts
$0.323
Median candidate latency
3.1 minutes

Model source: OpenRouter catalogue

Attributed costs include every OpenRouter-reported candidate and judge attempt for this model, including retries. The run total is authoritative for settled key spend.

OpenRouter provenance

Requested model and returned route.

Requested model
deepseek/deepseek-v4-pro
Routed provider
StreamLake
Endpoint tag
streamlake/fp8
Returned model
deepseek/deepseek-v4-pro
Quantisation
FP8
Pinned route price
streamlake/fp8: $0.713 input / $1.427 output per 1M tokens

Run 2026-07-18-final / 2026-07-18 / Closed-book

Dimension shape

Where the observed score comes from.

Overall combines deterministic task checks with blinded, calibrated, cross-family judgements of the IC note.

Grounding82.0
Commercial judgement90.7
Scepticism87.2
Numerical sanity100.0
Risk discovery81.1
Question generation91.3
European context85.0
Output usefulness91.5

Complete field

Position among ranked two-sample runs.

6498

Interpretation limit

One case, 2 scored samples.

Directional mini benchmark: one fresh synthetic case and two samples per model. It is not a universal ranking of model intelligence.