Model rankings
Find a model for your workload →Compare models, one benchmark at a time. Updated
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek V4.1 FlashDeepSeek | 74.2%Source ↗ |
| 2 | GPT-5.6 SolOpenAI | 72.7%Source ↗ |
| 3 | Grok 4.7SpaceXAI | 71.0%Source ↗ |
| 4 | Step 5 PreviewStepFun | 67.7%Source ↗ |
| 5 | GLM-5.3Z.ai | 66.9%Source ↗ |
| 6 | Grok 4.6xAI | 65.9%Source ↗ |
| 7 | Gemini 3.7 FlashGoogle | 65.3%Source ↗ |
| 8 | Kimi K3Moonshot AI | 63.8%Source ↗ |
| 9 | GLM-5.3-FlashZ.ai | 63.4%Source ↗ |
| 10 | GPT-5.4 miniOpenAI | 62.1%Source ↗ |
| 11 | MiniMax M3MiniMax | 61.2%Source ↗ |
| 12 | Qwen3.8-Flash-NextAlibaba | 58.7%Source ↗ |
| 13 | Qwen3.8-2.4T-A95BAlibaba | 56.6%Source ↗ |
Higher scores rank first. Publisher test setups vary. Prices blend input and output 3:1; missing results are excluded.
Compare score against price
Score vs price
DeepSWE v1.1 against current blended list price, 3:1 input to output. Up and left is better.
- DeepSeek V4.1 Flash: 74.2%, $0.26 / 1M (Off-peak)
- GPT-5.6 Sol: 72.7%, $11 / 1M (Current list price)
- Grok 4.7: 71.0%, $3 / 1M (Publisher starting price)
- Step 5 Preview: 67.7%, $1.43 / 1M (Vercel provider rates · checked Sep 21)
- Grok 4.6: 65.9%, $3 / 1M (Standard)
- Gemini 3.7 Flash: 65.3%, $1.5 / 1M (Introductory)
- Kimi K3: 63.8%, $1.05 / 1M (Standard)
- GLM-5.3-Flash: 63.4%, $0.24 / 1M (List)
- GPT-5.4 mini: 62.1%, $0.79 / 1M (Standard)
- MiniMax M3: 61.2%, $0.7 / 1M (Standard)
- Qwen3.8-2.4T-A95B: 56.6%, $2.48 / 1M (Model Studio API)
Scroll to compare the full chart →
Output tokens per second · Higher is better.
Source: Artificial Analysis ↗ · 2026-09-21. These are not MiniRouter route measurements.
One variant per family, selected by intelligence score. Prices blend input and output 3:1.
DeepSWE price performance
Reported score per $1M tokens at each published price, blended 3:1 input to output. Directional only.
| # | Model / price tier | Score / $ |
|---|---|---|
| 1 | GLM-5.3-FlashLaunch discount | 533.9$0.12 / 1M |
| 2 | DeepSeek V4.1 FlashOff-peak | 282.7$0.26 / 1M |
| 3 | GLM-5.3-FlashList | 266.9$0.24 / 1M |
| 4 | DeepSeek V4.1 FlashPeak · weekdays 01–04 / 06–10 UTC | 141.3$0.52 / 1M |
| 5 | MiniMax M3Standard | 87.4$0.70 / 1M |
| 6 | GPT-5.4 miniStandard | 78.9$0.79 / 1M |
| 7 | Kimi K3Standard | 60.8$1.05 / 1M |
| 8 | Step 5 PreviewVercel provider rates · checked Sep 21 | 47.5$1.43 / 1M |
| 9 | Gemini 3.7 FlashIntroductory | 43.5$1.50 / 1M |
| 10 | Grok 4.7Publisher starting price | 23.7$3.00 / 1M |
| 11 | Qwen3.8-2.4T-A95BModel Studio API | 22.9$2.48 / 1M |
| 12 | Grok 4.6Standard | 22.0$3.00 / 1M |
| 13 | Gemini 3.7 FlashStandard | 21.8$3.00 / 1M |
| 14 | GPT-5.6 SolCurrent list price | 6.5$11.25 / 1M |
Unpriced: GLM-5.3 66.9%, Qwen3.8-Flash-Next 58.7
Method, sources, and raw data
- What is ranked
- Publisher scores inside one named benchmark. No composite score.
- What a gap means
- No publisher result. MiniRouter does not estimate one.
- How to compare
- Check harness, scaffold, and effort settings before comparing.