Alibaba Announced Aug 26, 2026

Qwen3.8-Flash-Next

Alibaba's open-weight 125B MoE with 6B active parameters, positioned as an early preview of the Qwen4 architecture, with vision input, selectable reasoning effort, and publisher-reported coding and reasoning results.

Architecture
125B MoE · 6B active
Hybrid Gated DeltaNet and sparse attention, plus a 51B n-gram table
Context
262K native
Extensible to 1M with YaRN
Weights
Released
Hugging Face; runs on vLLM, SGLang, llama.cpp, and MLX
Reasoning effort
low / medium / xhigh
Thinking on by default; can be disabled per request

Publisher pricing

Pricing below comes from the release announcement and does not represent a MiniRouter route or price.

Not published

The publisher announcement does not list an API price. MiniRouter does not estimate one.

Publisher-reported benchmarks

Results are grouped by unit. Percentage, Elo, and task-count results never share a scale.

BenchmarkReported resultReported byAs ofSource
DeepSWE 1.158.7AlibabaAug 26, 2026Announcement ↗
SWE-Bench Pro62.5AlibabaAug 26, 2026Announcement ↗
GPQA Diamond91.7AlibabaAug 26, 2026Announcement ↗
SWE-Bench Multilingual81.0AlibabaAug 26, 2026Announcement ↗
CoWorkBench73.9AlibabaAug 26, 2026Announcement ↗
Agents' Last Exam51.2AlibabaAug 26, 2026Announcement ↗
LiveCodeBench v691.9AlibabaAug 26, 2026Announcement ↗
  • Alibaba says training cost about one-ninth of Qwen3.7-Plus while delivering stronger coding and office-task results.
  • Alibaba publishes no API price for the open checkpoint; the routable price on MiniRouter is the live route rate.
  • Alibaba reports the benchmark figures; MiniRouter has not reproduced them.

Independent measurements

Artificial Analysis runs its own evaluations and timing against each model's first-party API. These are the only numbers on this page not reported by the publisher.

Qwen3.8-Flash-Next ranks #17 of 448 model families on their Intelligence Index.

Intelligence Index
55.8
Composite of their evaluations
Coding Index
73.1
Coding evaluations only
Blended price
$0.23
USD per 1M tokens, 3:1 input to output, first-party list
Output speed
84
Median tokens per second
First answer token
25s
Median seconds, including reasoning

Qwen3.8-Flash-Next on Artificial Analysis ↗

Source: Artificial Analysis, read Sep 2, 2026. Measured on the model's first-party API, not a MiniRouter route.

What is not independently confirmed

  • Every figure on this page is self-reported by the cited publisher. MiniRouter did not run these benchmarks.
  • Benchmark harnesses are not standardised. Results sharing a name across vendors have not been confirmed to use identical runs.
  • Publisher pricing can change, and introductory phases have stated end dates.
  • Qwen3.8-Flash-Next is routable through MiniRouter today; live pricing on its model page is authoritative over the publisher figures here.

Related coverage

Other release dossiers and the current MiniRouter catalog.

Sources

  1. Qwen3.8-Flash-Next release notes

    Published Aug 26, 2026

    Read at Alibaba
  2. Artificial Analysis model page

    Published Sep 2, 2026

    Read at Artificial Analysis

Editorial reference revised 2026-09-02. Catalog status comes from the generated routing snapshot.