Z.ai Announced Aug 26, 2026

GLM-5.3-Flash

Z.ai's natively multimodal coding model, tested anonymously as ox-alpha, pairing 320B total with 18B active parameters, MIT-licensed weights, and a 50% launch discount through Sep 9, 2026.

Architecture
320B MoE · 18B active
Hybrid attention; the first natively multimodal GLM-5 model
Context
1M
128K maximum output per Z.ai's docs
Weights
Released
MIT license on Hugging Face
Reasoning effort
low / high / max
Z.ai recommends max for coding

Publisher pricing

Pricing below comes from the release announcement and does not represent a MiniRouter route or price.

Publisher-listed API price phases in US dollars per million tokens.
PhaseInput / 1MOutput / 1MBlended / 1MEffectiveSource
Launch discount$0.08$0.25$0.12Aug 26, 2026Sep 9, 2026Z.ai
List$0.15$0.50$0.24Sep 10, 2026open-endedZ.ai

Publisher-reported benchmarks

Results are grouped by unit. Percentage, Elo, and task-count results never share a scale.

BenchmarkReported resultReported byAs ofSource
Terminal Bench 2.184.3Z.aiAug 26, 2026Announcement ↗
DeepSWE v1.163.4Z.aiAug 26, 2026Announcement ↗
Agents' Last Exam26.3Z.aiAug 26, 2026Announcement ↗
AutomationBench v1.0.648.8Z.aiAug 26, 2026Announcement ↗
Humanity's Last Exam · with tools55.3Z.aiAug 26, 2026Announcement ↗
BenchmarkReported resultReported byAs ofSource
GDPval-AA v21,773Z.aiAug 26, 2026Announcement ↗
  • Z.ai says it tested the model anonymously as ox-alpha on OpenCode and OpenRouter before launch.
  • Z.ai published its benchmark charts as images; the values here are MiniRouter's transcription of those charts.
  • Z.ai reports the benchmark figures; MiniRouter has not reproduced them.

Independent measurements

Artificial Analysis runs its own evaluations and timing against each model's first-party API. These are the only numbers on this page not reported by the publisher.

GLM-5.3-Flash ranks #11 of 448 model families on their Intelligence Index.

Intelligence Index
57.5
Composite of their evaluations
Coding Index
71.5
Coding evaluations only
Blended price
$0.24
USD per 1M tokens, 3:1 input to output, first-party list
Output speed
44
Median tokens per second
First answer token
47s
Median seconds, including reasoning

GLM-5.3-Flash on Artificial Analysis ↗

Source: Artificial Analysis, read Sep 2, 2026. Measured on the model's first-party API, not a MiniRouter route.

What is not independently confirmed

  • Every figure on this page is self-reported by the cited publisher. MiniRouter did not run these benchmarks.
  • Benchmark harnesses are not standardised. Results sharing a name across vendors have not been confirmed to use identical runs.
  • Publisher pricing can change, and introductory phases have stated end dates.
  • GLM-5.3-Flash is routable through MiniRouter today; live pricing on its model page is authoritative over the publisher figures here.

Related coverage

Other release dossiers and the current MiniRouter catalog.

Sources

  1. Z.ai announcement

    Published Aug 26, 2026

    Read at Z.ai
  2. Artificial Analysis model page

    Published Sep 2, 2026

    Read at Artificial Analysis

Editorial reference revised 2026-09-02. Catalog status comes from the generated routing snapshot.