Coding model rankings
Pick a model for the work
Compare publisher-reported coding results by benchmark. No composite score. Missing results stay missing.
25 models22 routableRevised 2026-09-29
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek V4.1 FlashDeepSeek | 74.2%Source ↗ |
| 2 | GPT-5.6 SolOpenAI | 72.7%Source ↗ |
| 3 | Grok 4.7SpaceXAI | 71.0%Source ↗ |
| 4 | Step 5 PreviewStepFun | 67.7%Source ↗ |
| 5 | GLM-5.3Z.ai | 66.9%Source ↗ |
| 6 | Grok 4.6xAI | 65.9%Source ↗ |
| 7 | Gemini 3.7 FlashGoogle | 65.3%Source ↗ |
| 8 | Kimi K3Moonshot AI | 63.8%Source ↗ |
| 9 | GLM-5.3-FlashZ.ai | 63.4%Source ↗ |
| 10 | GPT-5.4 miniOpenAI | 62.1%Source ↗ |
| 11 | MiniMax M3MiniMax | 61.2%Source ↗ |
| 12 | Qwen3.8-Flash-NextAlibaba | 58.7%Source ↗ |
| 13 | Qwen3.8-2.4T-A95BAlibaba | 56.6%Source ↗ |
Higher scores rank first. Publisher test setups vary. Prices blend input and output 3:1; missing results are excluded.
Compare score against price
Scroll to compare the full chart →
Method, sources, and raw data
What is ranked
Publisher-reported scores. Compare within each benchmark.
What a gap means
The publisher did not report that result. MiniRouter does not estimate it or turn it into zero.
How to compare
Harnesses, scaffolds, and effort settings may differ even when vendors use the same benchmark name.