Alibaba

Qwen 3.8 Max Prime

Qwen 3.8 Max served faster. Same 2.4T model, 1.5 to 2x the output speed, twice the price.

qwen3.8-max, high-speed editionText + visiontext + image inputtext output

Run this model

alibaba/qwen3.8-max-prime

Use an endpoint supported by this model.

Get a key for this setup →
const response = await fetch('https://api.minirouter.sh/v1/chat/completions', {  method: 'POST',  headers: {    Authorization: `Bearer ${process.env.MINIROUTER_KEY}`,    'Content-Type': 'application/json',  },  body: JSON.stringify({    model: 'alibaba/qwen3.8-max-prime',    messages: [{ role: 'user', content: 'Why is the sky blue?' }],    max_tokens: 1024,  }),}) const data = await response.json()console.log(data.choices[0].message.content)
Read docs →

Model overview

Input / 1M
$4.2
All-in MiniRouter rate
Output / 1M
$12.6
All-in MiniRouter rate
Context
1M
1,000,000 tokens
Max output
131K
tokens
Speed
1.5 to 2x
Alibaba's claim versus Qwen 3.8 Max
Price
2x Qwen 3.8 Max
$4 in, $12 out per 1M
Released
23 Sept 2026
publisher release

Publisher-reported, plotted by MiniRouter

Max's scores, served faster.

Prime serves the Qwen 3.8 Max model, so these are Max's results from Qwen's release post. MiniRouter did not run these evaluations.

Qwen-reported

Coding, research and tools

  • Qwen 3.8 Max
  • Comparison models

Terminal

Terminal Bench 2.1

  • Qwen 3.8 Max86.6%
  • Claude Opus 4.884.6%
  • Claude Fable 584.6%
  • GPT-5.6 Sol88.8%
  • Qwen 3.7 Max74.5%

0 to 100%

Coding agent

SWE-bench Pro

  • Qwen 3.8 Max67.7%
  • Claude Opus 4.869.2%
  • Claude Fable 580.0%
  • GPT-5.6 Sol64.6%
  • Qwen 3.7 Max60.6%

0 to 100%

Research

PaperBench

  • Qwen 3.8 Max93.0%
  • Claude Opus 4.880.3%
  • Claude Fable 588.8%
  • GPT-5.6 Sol90.5%
  • Qwen 3.7 Max64.8%

0 to 100%

Tool use

Toolathlon Verified

  • Qwen 3.8 Max72.5%
  • Claude Opus 4.876.2%
  • Claude Fable 577.9%
  • GPT-5.6 Sol74.9%
  • Qwen 3.7 Max49.7%

0 to 100%

Qwen's harness and settings; comparison scores as Qwen prints them, some taken from other publishers. Source: Qwen3.8-Max release post ↗, August 2026.
View chart data
Values from Qwen's Qwen3.8-Max release post
BenchmarkQwen 3.8 MaxClaude Opus 4.8Claude Fable 5GPT-5.6 SolQwen 3.7 Max
Terminal Bench 2.186.6%84.6%84.6%88.8%74.5%
SWE-bench Pro67.7%69.2%80.0%64.6%60.6%
PaperBench93.0%80.3%88.8%90.5%64.8%
Toolathlon Verified72.5%76.2%77.9%74.9%49.7%

When it pays off

Pay for speed only where you wait.

Route interactive turns here and background work to Qwen 3.8 Max.

Qwen 3.8 Max

Vercel list, per 1M tokens

Input
$4.00$2.00 standard
Output
$12.00$6.00 standard
Cache read
$0.50$0.25 standard
Source: Vercel AI Gateway model page ↗, September 24, 2026.
View chart data
Vercel list prices per 1M tokens
RateMax PrimeMax
Input$4.00$2.00
Output$12.00$6.00
Cache read$0.50$0.25

One route, standard clients.

Authenticate with a MiniRouter bearer key. Caller-provided upstream credentials are not used on this route.

Read model setup docs →
Base URLhttps://api.minirouter.sh/v1
Model IDalibaba/qwen3.8-max-prime
Endpoints
/v1/chat/completions/v1/messages/v1/responses
Parameters

max_tokens, temperature, stop, tools, tool_choice, reasoning, include_reasoning, response_format, structured_outputs

Structured output

JSON Schema and JSON mode enforced on verified providers.

Configured provider

Priority follows the live catalog order.

Input: text, image · Output: text

ProviderPriorityInput / 1MOutput / 1MParametersAccess
vercel
1$4.2$12.6max_tokens, temperature, stop, tools, tool_choice, reasoning, include_reasoning, response_format, structured_outputsMiniRouter balance