Qwen3.8-Flash-Next
Alibaba's open-weight 125B MoE with 6B active parameters, positioned as an early preview of the Qwen4 architecture, with vision input, selectable reasoning effort, and publisher-reported coding and reasoning results.
Routable on MiniRouter
Use alibaba/qwen3.8-flash-next via the API →- Architecture
- 125B MoE · 6B active
- Hybrid Gated DeltaNet and sparse attention, plus a 51B n-gram table
- Context
- 262K native
- Extensible to 1M with YaRN
- Weights
- Released
- Hugging Face; runs on vLLM, SGLang, llama.cpp, and MLX
- Reasoning effort
- low / medium / xhigh
- Thinking on by default; can be disabled per request
Publisher pricing
Pricing below comes from the release announcement and does not represent a MiniRouter route or price.
Not published
The publisher announcement does not list an API price. MiniRouter does not estimate one.
Publisher-reported benchmarks
Results are grouped by unit. Percentage, Elo, and task-count results never share a scale.
| Benchmark | Reported result | Reported by | As of | Source |
|---|---|---|---|---|
| DeepSWE 1.1 | 58.7 | Alibaba | Aug 26, 2026 | Announcement ↗ |
| SWE-Bench Pro | 62.5 | Alibaba | Aug 26, 2026 | Announcement ↗ |
| GPQA Diamond | 91.7 | Alibaba | Aug 26, 2026 | Announcement ↗ |
| SWE-Bench Multilingual | 81.0 | Alibaba | Aug 26, 2026 | Announcement ↗ |
| CoWorkBench | 73.9 | Alibaba | Aug 26, 2026 | Announcement ↗ |
| Agents' Last Exam | 51.2 | Alibaba | Aug 26, 2026 | Announcement ↗ |
| LiveCodeBench v6 | 91.9 | Alibaba | Aug 26, 2026 | Announcement ↗ |
- — Alibaba says training cost about one-ninth of Qwen3.7-Plus while delivering stronger coding and office-task results.
- — Alibaba publishes no API price for the open checkpoint; the routable price on MiniRouter is the live route rate.
- — Alibaba reports the benchmark figures; MiniRouter has not reproduced them.
Independent measurements
Artificial Analysis runs its own evaluations and timing against each model's first-party API. These are the only numbers on this page not reported by the publisher.
Qwen3.8-Flash-Next ranks #17 of 448 model families on their Intelligence Index.
- Intelligence Index
- 55.8
- Composite of their evaluations
- Coding Index
- 73.1
- Coding evaluations only
- Blended price
- $0.23
- USD per 1M tokens, 3:1 input to output, first-party list
- Output speed
- 84
- Median tokens per second
- First answer token
- 25s
- Median seconds, including reasoning
Qwen3.8-Flash-Next on Artificial Analysis ↗
Source: Artificial Analysis ↗, read Sep 2, 2026. Measured on the model's first-party API, not a MiniRouter route.
What is not independently confirmed
- Every figure on this page is self-reported by the cited publisher. MiniRouter did not run these benchmarks.
- Benchmark harnesses are not standardised. Results sharing a name across vendors have not been confirmed to use identical runs.
- Publisher pricing can change, and introductory phases have stated end dates.
- Qwen3.8-Flash-Next is routable through MiniRouter today; live pricing on its model page is authoritative over the publisher figures here.
Related coverage
Other release dossiers and the current MiniRouter catalog.
Sources
- Read at Alibaba ↗
Qwen3.8-Flash-Next release notes
Published Aug 26, 2026
- Read at Artificial Analysis ↗
Artificial Analysis model page
Published Sep 2, 2026
Editorial reference revised 2026-09-02. Catalog status comes from the generated routing snapshot.