MidMarketBench
Rank 4completeObserved OpenRouter runProvisional tie

Gemini 3.5 Flash

Model and run

Provider
Google
Version
gemini-3.5-flash
Release
Not published
Context
1,048,576 tokens
Weights
Not published
Run status
complete
Candidate samples
2/2
Scored samples
2/2
Overall
89.1
Observed range
86.9 to 91.3
Provider-reported attributed cost
$0.725
Candidate attempts
$0.099
Judge attempts
$0.627
Median candidate latency
2.0 minutes

Model source: OpenRouter catalogue

Attributed costs include every OpenRouter-reported candidate and judge attempt for this model, including retries. The run total is authoritative for settled key spend.

OpenRouter provenance

Requested model and returned route.

Requested model
google/gemini-3.5-flash
Routed provider
Google
Endpoint tag
google-vertex/global
Returned model
google/gemini-3.5-flash
Quantisation
Not reported
Pinned route price
google-vertex/global: $1.500 input / $9.000 output per 1M tokens

Run 2026-07-18-final / 2026-07-18 / Closed-book

Dimension shape

Where the observed score comes from.

Overall combines deterministic task checks with blinded, calibrated, cross-family judgements of the IC note.

Grounding86.7
Commercial judgement81.1
Scepticism88.8
Numerical sanity100.0
Risk discovery91.7
Question generation92.1
European context96.7
Output usefulness74.5

Complete field

Position among ranked two-sample runs.

6498

Interpretation limit

One case, 2 scored samples.

Directional mini benchmark: one fresh synthetic case and two samples per model. It is not a universal ranking of model intelligence.