GLM-5.3-Flash
Z.ai's natively multimodal coding model, tested anonymously as ox-alpha, pairing 320B total with 18B active parameters, MIT-licensed weights, and a 50% launch discount through Sep 9, 2026.
Routable on MiniRouter
Use zai/glm-5.3-flash via the API →- Architecture
- 320B MoE · 18B active
- Hybrid attention; the first natively multimodal GLM-5 model
- Context
- 1M
- 128K maximum output per Z.ai's docs
- Weights
- Released
- MIT license on Hugging Face
- Reasoning effort
- low / high / max
- Z.ai recommends max for coding
Publisher pricing
Pricing below comes from the release announcement and does not represent a MiniRouter route or price.
Publisher-reported benchmarks
Results are grouped by unit. Percentage, Elo, and task-count results never share a scale.
| Benchmark | Reported result | Reported by | As of | Source |
|---|---|---|---|---|
| Terminal Bench 2.1 | 84.3 | Z.ai | Aug 26, 2026 | Announcement ↗ |
| DeepSWE v1.1 | 63.4 | Z.ai | Aug 26, 2026 | Announcement ↗ |
| Agents' Last Exam | 26.3 | Z.ai | Aug 26, 2026 | Announcement ↗ |
| AutomationBench v1.0.6 | 48.8 | Z.ai | Aug 26, 2026 | Announcement ↗ |
| Humanity's Last Exam · with tools | 55.3 | Z.ai | Aug 26, 2026 | Announcement ↗ |
| Benchmark | Reported result | Reported by | As of | Source |
|---|---|---|---|---|
| GDPval-AA v2 | 1,773 | Z.ai | Aug 26, 2026 | Announcement ↗ |
- — Z.ai says it tested the model anonymously as ox-alpha on OpenCode and OpenRouter before launch.
- — Z.ai published its benchmark charts as images; the values here are MiniRouter's transcription of those charts.
- — Z.ai reports the benchmark figures; MiniRouter has not reproduced them.
Independent measurements
Artificial Analysis runs its own evaluations and timing against each model's first-party API. These are the only numbers on this page not reported by the publisher.
GLM-5.3-Flash ranks #11 of 448 model families on their Intelligence Index.
- Intelligence Index
- 57.5
- Composite of their evaluations
- Coding Index
- 71.5
- Coding evaluations only
- Blended price
- $0.24
- USD per 1M tokens, 3:1 input to output, first-party list
- Output speed
- 44
- Median tokens per second
- First answer token
- 47s
- Median seconds, including reasoning
GLM-5.3-Flash on Artificial Analysis ↗
Source: Artificial Analysis ↗, read Sep 2, 2026. Measured on the model's first-party API, not a MiniRouter route.
What is not independently confirmed
- Every figure on this page is self-reported by the cited publisher. MiniRouter did not run these benchmarks.
- Benchmark harnesses are not standardised. Results sharing a name across vendors have not been confirmed to use identical runs.
- Publisher pricing can change, and introductory phases have stated end dates.
- GLM-5.3-Flash is routable through MiniRouter today; live pricing on its model page is authoritative over the publisher figures here.
Related coverage
Other release dossiers and the current MiniRouter catalog.
Sources
- Read at Z.ai ↗
Z.ai announcement
Published Aug 26, 2026
- Read at Artificial Analysis ↗
Artificial Analysis model page
Published Sep 2, 2026
Editorial reference revised 2026-09-02. Catalog status comes from the generated routing snapshot.