Calculate eleven operating, retention, margin and market metrics with formulae and exact evidence IDs.
Methodology v0.4-mini
Judgement carries the most weight.
The 18 July 2026 mini benchmark scores observed OpenRouter outputs from a fresh synthetic diligence packet. Overall is derived from eight dimensions and two samples per available model.
Dimensions and weights
| Dimension | Weight | Definition | High score | Low score |
|---|---|---|---|---|
| Grounding | 15% | Claims are supported by packet evidence or clearly marked as inference. | Specific evidence, disciplined attribution and no invented facts. | Generic claims, hallucinated facts or management language repeated as truth. |
| Commercial judgementHeaviest | 20% | Identifies what matters economically and strategically to the deal. | Prioritises revenue quality, durability and decision impact. | Treats all issues equally and misses the business model's economic engine. |
| Scepticism | 15% | Challenges weak claims, inflated metrics and convenient narratives. | Tests TAM, retention, ARR definitions and internal contradictions. | Accepts the CIM framing and assumes favourable interpretations. |
| Numerical sanity | 15% | Interprets and reconciles financial and operating data correctly. | Connects cohorts, concentration, ARR bridges and revenue mix. | Confuses metrics, ignores tables or makes arithmetic errors. |
| Risk discovery | 15% | Finds hidden and second-order risks across the packet. | Connects evidence across files and ranks risks by severity. | Lists generic risks and misses disguised business model weakness. |
| Question generation | 10% | Asks the next questions most likely to change the recommendation. | Specific, answerable, prioritised and disconfirmatory questions. | Long, generic diligence checklists with no decision hierarchy. |
| European context | 5% | Accounts for fragmented markets, regulation, language and procurement. | Avoids US category assumptions and tests country-level repeatability. | Treats Europe as one market and overstates reachable scale. |
| Output usefulness | 5% | Produces material a deal team can reuse with limited editing. | Concise, structured, memo-ready and explicit about next actions. | Verbose, generic and operationally difficult to reuse. |
Scoring reference
Anchored 0 to 4 rubric
Blinded judges grade five subjective dimensions against the same strong and weak reference notes. Each grade is normalised per judge and dimension with the weak anchor at 0 and strong anchor at 100 before panel aggregation.
| 4 | Exceptional | Exact, decision-changing and unusually well calibrated to the evidence and task. |
| 3 | Strong | Grounded, commercially useful and clear, with only limited omissions. |
| 2 | Acceptable | Broadly sound but generic, incomplete or weakly prioritised. |
| 1 | Weak | Touches the issue without enough evidence or decision relevance. |
| 0 | Miss | Misses or contradicts the evidence, or produces unusable work. |
Observed run protocol
Four tasks, one fresh packet.
Return six company-specific risks, challenged claims, supporting evidence and linked diligence actions.
Select exactly four actions within EUR 25k and eight parallel working days.
Make the EUR 250k diligence decision, identify thesis breakers and state decision-changing conditions.
Scoring operation
Hybrid, traceable evaluation.
Deterministic checks score arithmetic, citations, issue coverage, risk ordering, action linkage, budget feasibility and schema compliance. Blinded judges score only the IC note for grounding, commercial judgement, scepticism, European context and output usefulness.
Judge controls
Calibrated and cross-family.
Each judge also grades strong and weak anchors. Any dimension without separation is rejected. A judge from the same provider family as the candidate is excluded, then a three-judge median or two-judge mean is used. Material disagreement is recorded for review.
Execution controls
Schema support is checked twice.
Routes must support the requested reasoning, maximum-token and structured-output parameters. Returned JSON is then validated locally. The run records two frozen response-format hashes: a bounded schema and a provider-compatible variant. Removed bounds are checked as scored protocol constraints. Response completeness and instruction-injection resistance also contribute to compliance, with injection-following outputs capped at 25 overall. One provider-compatible Opus sample returned a fifth action and was penalised rather than discarded, so schema adaptation remains a comparability limit.
Reasoning policy
Bounded where providers allow it.
Candidates receive a 2,048-token reasoning budget and judges receive 1,024. Kimi K3 requires maximum reasoning, so that exception is explicit in its provenance rather than presented as an identical setting.
Routing and spend
Pinned, inspectable runs.
The runner snapshots OpenRouter's model catalogue and eligible endpoints, disables fallback on selected routes, records returned provider metadata and maintains a conservative in-process spend guard. The provider-side key cap is the authoritative cross-process limit.
Ranking policy
Complete evidence only.
The leaderboard averages two scored samples. A gap below two overall points is marked as a provisional tie. Partial or unavailable models remain visible for provenance but are excluded from ranked positions.
Scope limit
Directional, not universal.
One synthetic company, four tasks and two samples reveal behaviour on this workflow. They do not establish a universal ranking of model intelligence, reliability or value across other domains.
Run qualification: the candidate boundary explicitly treated packet text as untrusted. The v0.4 judge system instruction did not restate that boundary, although no accepted judgement followed the embedded instruction. Some early length failures retain cost and error metadata but not raw response text. Both controls are hardened for subsequent runs without mixing protocols into this dated result.