MidMarketBench
Rank 5completeObserved OpenRouter runProvisional tie

Claude Opus 4.8

Model and run

Provider
Anthropic
Version
claude-opus-4.8
Release
Not published
Context
1,000,000 tokens
Weights
Not published
Run status
complete
Candidate samples
2/2
Scored samples
2/2
Overall
88.8
Observed range
80.2 to 97.4
Provider-reported attributed cost
$0.649
Candidate attempts
$0.518
Judge attempts
$0.130
Median candidate latency
1.3 minutes

Model source: OpenRouter catalogue

Attributed costs include every OpenRouter-reported candidate and judge attempt for this model, including retries. The run total is authoritative for settled key spend.

OpenRouter provenance

Requested model and returned route.

Requested model
anthropic/claude-opus-4.8
Routed provider
Azure
Endpoint tag
azure/us-east-2
Returned model
anthropic/claude-opus-4.8
Quantisation
Not reported
Pinned route price
azure/us-east-2: $5.000 input / $25.000 output per 1M tokens

Run 2026-07-18-final / 2026-07-18 / Closed-book

Dimension shape

Where the observed score comes from.

Overall combines deterministic task checks with blinded, calibrated, cross-family judgements of the IC note.

Grounding91.3
Commercial judgement77.7
Scepticism96.5
Numerical sanity100.0
Risk discovery94.2
Question generation66.5
European context100.0
Output usefulness85.7

Complete field

Position among ranked two-sample runs.

6498

Interpretation limit

One case, 2 scored samples.

Directional mini benchmark: one fresh synthetic case and two samples per model. It is not a universal ranking of model intelligence.