MidMarketBench

Complete runs

Ranked on complete evidence.

RankModelProviderOverallRangeSamplesOpenRouter routeReported attributed costMedian latency
1Kimi K3Route caveatMoonshot AI96.395.6 to 96.92/2Moonshot AI / INT4$0.8946.4m
2Grok 4.5xAI96.395.5 to 97.02/2xAI$0.3281.5m
3GPT-5.6 SolOpenAI95.595.1 to 95.92/2OpenAI$0.5591.7m
4Gemini 3.5 FlashGoogle89.186.9 to 91.32/2Google$0.7252.0m
5Claude Opus 4.8Anthropic88.880.2 to 97.42/2Azure$0.6491.3m
6DeepSeek V4 ProDeepSeek88.687.9 to 89.42/2StreamLake / FP8$0.3613.1m
7GLM 5.2Z.ai88.286.6 to 89.72/2StreamLake / FP8$0.2612.3m
8Qwen3.7 MaxAlibaba65.544.6 to 86.42/2Alibaba$0.4111.3m

Not ranked

Unavailable or incomplete runs.

Infrastructure and routing failures are not converted into zero scores. These models remain in the run ledger without a rank.

ModelProviderStatusCandidate samplesOpenRouter routeRecorded reason
Claude Fable 5Anthropicunavailable0/2anthropic/claude-fable-5Claude Fable 5 is not available. Learn more: https://www.anthropic.com/news/fable-mythos-access