MidMarketBench

Dimensions and weights

DimensionWeightDefinitionHigh scoreLow score
Grounding15%Claims are supported by packet evidence or clearly marked as inference.Specific evidence, disciplined attribution and no invented facts.Generic claims, hallucinated facts or management language repeated as truth.
Commercial judgementHeaviest20%Identifies what matters economically and strategically to the deal.Prioritises revenue quality, durability and decision impact.Treats all issues equally and misses the business model's economic engine.
Scepticism15%Challenges weak claims, inflated metrics and convenient narratives.Tests TAM, retention, ARR definitions and internal contradictions.Accepts the CIM framing and assumes favourable interpretations.
Numerical sanity15%Interprets and reconciles financial and operating data correctly.Connects cohorts, concentration, ARR bridges and revenue mix.Confuses metrics, ignores tables or makes arithmetic errors.
Risk discovery15%Finds hidden and second-order risks across the packet.Connects evidence across files and ranks risks by severity.Lists generic risks and misses disguised business model weakness.
Question generation10%Asks the next questions most likely to change the recommendation.Specific, answerable, prioritised and disconfirmatory questions.Long, generic diligence checklists with no decision hierarchy.
European context5%Accounts for fragmented markets, regulation, language and procurement.Avoids US category assumptions and tests country-level repeatability.Treats Europe as one market and overstates reachable scale.
Output usefulness5%Produces material a deal team can reuse with limited editing.Concise, structured, memo-ready and explicit about next actions.Verbose, generic and operationally difficult to reuse.

Scoring reference

Anchored 0 to 4 rubric

Blinded judges grade five subjective dimensions against the same strong and weak reference notes. Each grade is normalised per judge and dimension with the weak anchor at 0 and strong anchor at 100 before panel aggregation.

4ExceptionalExact, decision-changing and unusually well calibrated to the evidence and task.
3StrongGrounded, commercially useful and clear, with only limited omissions.
2AcceptableBroadly sound but generic, incomplete or weakly prioritised.
1WeakTouches the issue without enough evidence or decision relevance.
0MissMisses or contradicts the evidence, or produces unusable work.

Observed run protocol

Four tasks, one fresh packet.

Metric reconstruction

Calculate eleven operating, retention, margin and market metrics with formulae and exact evidence IDs.

Ranked red flags

Return six company-specific risks, challenged claims, supporting evidence and linked diligence actions.

Bounded diligence

Select exactly four actions within EUR 25k and eight parallel working days.

IC decision note

Make the EUR 250k diligence decision, identify thesis breakers and state decision-changing conditions.

Scoring operation

Hybrid, traceable evaluation.

Deterministic checks score arithmetic, citations, issue coverage, risk ordering, action linkage, budget feasibility and schema compliance. Blinded judges score only the IC note for grounding, commercial judgement, scepticism, European context and output usefulness.

Judge controls

Calibrated and cross-family.

Each judge also grades strong and weak anchors. Any dimension without separation is rejected. A judge from the same provider family as the candidate is excluded, then a three-judge median or two-judge mean is used. Material disagreement is recorded for review.

Execution controls

Schema support is checked twice.

Routes must support the requested reasoning, maximum-token and structured-output parameters. Returned JSON is then validated locally. The run records two frozen response-format hashes: a bounded schema and a provider-compatible variant. Removed bounds are checked as scored protocol constraints. Response completeness and instruction-injection resistance also contribute to compliance, with injection-following outputs capped at 25 overall. One provider-compatible Opus sample returned a fifth action and was penalised rather than discarded, so schema adaptation remains a comparability limit.

Reasoning policy

Bounded where providers allow it.

Candidates receive a 2,048-token reasoning budget and judges receive 1,024. Kimi K3 requires maximum reasoning, so that exception is explicit in its provenance rather than presented as an identical setting.

Routing and spend

Pinned, inspectable runs.

The runner snapshots OpenRouter's model catalogue and eligible endpoints, disables fallback on selected routes, records returned provider metadata and maintains a conservative in-process spend guard. The provider-side key cap is the authoritative cross-process limit.

Ranking policy

Complete evidence only.

The leaderboard averages two scored samples. A gap below two overall points is marked as a provisional tie. Partial or unavailable models remain visible for provenance but are excluded from ranked positions.

Scope limit

Directional, not universal.

One synthetic company, four tasks and two samples reveal behaviour on this workflow. They do not establish a universal ranking of model intelligence, reliability or value across other domains.

Run qualification: the candidate boundary explicitly treated packet text as untrusted. The v0.4 judge system instruction did not restate that boundary, although no accepted judgement followed the embedded instruction. Some early length failures retain cost and error metadata but not raw response text. Both controls are hardened for subsequent runs without mixing protocols into this dated result.