Model rankings

Find a model for your workload →

Compare models, one benchmark at a time. Updated

DeepSWE v1.1

13 models · publisher-reported scores

Higher scores rank first. Publisher test setups vary. Prices blend input and output 3:1; missing results are excluded.

Compare score against price

Score vs price

DeepSWE v1.1 against current blended list price, 3:1 input to output. Up and left is better.

50556065707580$0.2$0.5$1$2$5$10$20$ per 1M tokens, log scaleDeepSeek V4.1 FlashGPT-5.6 SolGrok 4.7Step 5 PreviewGrok 4.6Gemini 3.7 FlashKimi K3GLM-5.3-FlashGPT-5.4 miniMiniMax M3Qwen3.8-2.4T-A95B

Scroll to compare the full chart →

Method, sources, and raw data
What is ranked
Publisher scores inside one named benchmark. No composite score.
What a gap means
No publisher result. MiniRouter does not estimate one.
How to compare
Check harness, scaffold, and effort settings before comparing.