Coding model reference

Coding benchmarks, as reported by publishers

Every figure below is self-reported by the model's publisher. Harnesses can differ even where a benchmark name matches, so this is not an independent or like-for-like ranking.

Price-performance field

DeepSWE v1.1 scores against a 3:1 input-output blended price. The two publisher runs may use different harnesses, so position is orientation—not a rank.

Price as published · Aug 14, 2026

DeepSWE v1.1 publisher-reported scores against blended API price. Gemini 3.7 Flash is 65.3 percent at 1 dollar 50 introductory and 3 dollars standard per million blended tokens. GLM-5.3 is 66.9 percent with no published price.

Price not published

GLM-5.3

Blended price = 0.75 × input price + 0.25 × output price. Hollow squares mean not in MiniRouter's catalog; GLM-5.3 remains on the unpriced rail instead of receiving an estimate.

Scroll for more columns →

Visible table equivalent of the field. Publisher-reported results; no overall rank is assigned.
ModelCatalogPrice phaseBlended / 1MDeepSWE v1.1Reported bySource
Gemini 3.7 FlashNoIntroductory$1.5065.3%GoogleGoogle ↗
Gemini 3.7 FlashNoStandard$3.0065.3%GoogleGoogle ↗
GLM-5.3NoNot published66.9%Z.aiZ.ai ↗

Full reported benchmark inventory

A dash means the publisher announcement did not report that result. It does not mean zero.

Percentage results

Scroll for more columns →

Percentage results. Each value is publisher-reported and links to its source.
BenchmarkGemini 3.7 FlashGLM-5.3DeepSeek-V4-ProGrok 4.6Muse Glimmer 30BQwen3.8-2.4T-A95B
FrontierCode 1.143.6%Google · 2026-08-13
DeepSWE v1.165.3%Google · 2026-08-1366.9%Z.ai · 2026-08-1465.9%xAI · 2026-08-1256.6Alibaba Cloud · 2026-08-03
Terminal Bench 3.028.3%Z.ai · 2026-08-1426%xAI · 2026-08-12
CyberGym84.5%Z.ai · 2026-08-1483.3DeepSeek · 2026-08-13
ExploitBench54.4%Z.ai · 2026-08-14
Terminal Bench 2.187.9DeepSeek · 2026-08-1351.7Meta · 2026-08-1086.6Alibaba Cloud · 2026-08-03
HLE (with tools)60.0DeepSeek · 2026-08-13
DeepSWE62.7DeepSeek · 2026-08-13
Toolathlon-Verified74.1DeepSeek · 2026-08-13
FrontierCode v1.1 (Extended)61.3%xAI · 2026-08-12
CursorBench v3.269.9%xAI · 2026-08-12
APEX-SWE56.4%xAI · 2026-08-12
SWE-Bench Verified76.0Meta · 2026-08-10
SWE-Bench Pro51.2Meta · 2026-08-1067.7Alibaba Cloud · 2026-08-03
OSWorld-Verified65.9Meta · 2026-08-10
GPQA Diamond83.5Meta · 2026-08-1092.6Alibaba Cloud · 2026-08-03
HLE43.6Alibaba Cloud · 2026-08-03

Elo results

Scroll for more columns →

Elo results. Each value is publisher-reported and links to its source.
BenchmarkGemini 3.7 FlashGLM-5.3DeepSeek-V4-ProGrok 4.6Muse Glimmer 30BQwen3.8-2.4T-A95B
WebDev Arena1588 EloGoogle · 2026-08-13
GDPVal-AA v21753xAI · 2026-08-12

Task-count results

Scroll for more columns →

Task-count results. Each value is publisher-reported and links to its source.
BenchmarkGemini 3.7 FlashGLM-5.3DeepSeek-V4-ProGrok 4.6Muse Glimmer 30BQwen3.8-2.4T-A95B
ExploitGym · 2 hours105 tasksZ.ai · 2026-08-14
ExploitGym · 6 hours130 tasksZ.ai · 2026-08-14

What is not independently confirmed

  • Every figure on this page is self-reported by the cited publisher. MiniRouter did not run these benchmarks.
  • Benchmark harnesses are not standardised. Results sharing a name across vendors have not been confirmed to use identical runs.
  • Publisher pricing can change, and introductory phases have stated end dates.
  • Catalog status is shown per model above. Dossier coverage is not a commitment to keep any route available.

Sources

  1. Google announcement

    Published Aug 13, 2026

    Read at Google
  2. Z.ai announcement

    Published Aug 14, 2026

    Read at Z.ai
  3. DeepSeek GA release notes

    Published Aug 13, 2026

    Read at DeepSeek
  4. xAI announcement

    Published Aug 12, 2026

    Read at xAI
  5. Meta AI Research announcement

    Published Aug 10, 2026

    Read at Meta
  6. Alibaba Cloud announcement

    Published Aug 3, 2026

    Read at Alibaba Cloud

Page revised 2026-08-14. Dates change only when the content changes.