What it measures
Grounded commercial judgement.
Standard prompts, reference-scored calculations, bounded diligence choices and blinded judgements of decision-useful IC work.
About
MidMarketBench evaluates observed frontier-model work against a fresh synthetic European private-market case.
What it measures
Standard prompts, reference-scored calculations, bounded diligence choices and blinded judgements of decision-useful IC work.
What it does not measure
Norwyn Controls is not a real company. Two runs against one case are too small to establish general model quality, and this leaderboard should not be read as investment advice.
Provenance
Norwyn Controls was created for the 18 July 2026 run from realistic operating patterns, without proprietary deal material. Candidate and judge responses were generated through OpenRouter. Each artefact records the requested model, routed endpoint, provider metadata, prompt and case hashes, token use, latency and cost.
Comparability
Each available candidate received four closed-book tasks twice. Candidates used a 2,048-token reasoning budget and judges used 1,024; Kimi K3 ran at its provider-mandated maximum reasoning level.
Interpretation
Models within two overall points are marked as provisional ties. Models without a complete, valid scored run are shown as unavailable or partial and are excluded from the ranked field.
Audience
PE and growth investors evaluating model usefulness in diligence and origination.
Internal teams selecting models, prompts, tools and evaluation controls.
Vendors testing whether general capability translates into decision-useful work.