Alibaba
Qwen 3.8 Max Prime
Qwen 3.8 Max served faster. Same 2.4T model, 1.5 to 2x the output speed, twice the price.
Run this model
alibaba/qwen3.8-max-prime
Use an endpoint supported by this model.
Get a key for this setup →const response = await fetch('https://api.minirouter.sh/v1/chat/completions', { method: 'POST', headers: { Authorization: `Bearer ${process.env.MINIROUTER_KEY}`, 'Content-Type': 'application/json', }, body: JSON.stringify({ model: 'alibaba/qwen3.8-max-prime', messages: [{ role: 'user', content: 'Why is the sky blue?' }], max_tokens: 1024, }),}) const data = await response.json()console.log(data.choices[0].message.content)Model overview
- Input / 1M
- $4.2
- All-in MiniRouter rate
- Output / 1M
- $12.6
- All-in MiniRouter rate
- Context
- 1M
- 1,000,000 tokens
- Max output
- 131K
- tokens
- Speed
- 1.5 to 2x
- Alibaba's claim versus Qwen 3.8 Max
- Price
- 2x Qwen 3.8 Max
- $4 in, $12 out per 1M
- Released
- 23 Sept 2026
- publisher release
Publisher-reported, plotted by MiniRouter
Max's scores, served faster.
Prime serves the Qwen 3.8 Max model, so these are Max's results from Qwen's release post. MiniRouter did not run these evaluations.
Qwen-reported
Coding, research and tools
- Qwen 3.8 Max
- Comparison models
Terminal
Terminal Bench 2.1
- Qwen 3.8 Max86.6%
- Claude Opus 4.884.6%
- Claude Fable 584.6%
- GPT-5.6 Sol88.8%
- Qwen 3.7 Max74.5%
0 to 100%
Coding agent
SWE-bench Pro
- Qwen 3.8 Max67.7%
- Claude Opus 4.869.2%
- Claude Fable 580.0%
- GPT-5.6 Sol64.6%
- Qwen 3.7 Max60.6%
0 to 100%
Research
PaperBench
- Qwen 3.8 Max93.0%
- Claude Opus 4.880.3%
- Claude Fable 588.8%
- GPT-5.6 Sol90.5%
- Qwen 3.7 Max64.8%
0 to 100%
Tool use
Toolathlon Verified
- Qwen 3.8 Max72.5%
- Claude Opus 4.876.2%
- Claude Fable 577.9%
- GPT-5.6 Sol74.9%
- Qwen 3.7 Max49.7%
0 to 100%
View chart data
| Benchmark | Qwen 3.8 Max | Claude Opus 4.8 | Claude Fable 5 | GPT-5.6 Sol | Qwen 3.7 Max |
|---|---|---|---|---|---|
| Terminal Bench 2.1 | 86.6% | 84.6% | 84.6% | 88.8% | 74.5% |
| SWE-bench Pro | 67.7% | 69.2% | 80.0% | 64.6% | 60.6% |
| PaperBench | 93.0% | 80.3% | 88.8% | 90.5% | 64.8% |
| Toolathlon Verified | 72.5% | 76.2% | 77.9% | 74.9% | 49.7% |
When it pays off
Pay for speed only where you wait.
Route interactive turns here and background work to Qwen 3.8 Max.
Qwen 3.8 MaxVercel list, per 1M tokens
- Input
- $4.00$2.00 standard
- Output
- $12.00$6.00 standard
- Cache read
- $0.50$0.25 standard
View chart data
| Rate | Max Prime | Max |
|---|---|---|
| Input | $4.00 | $2.00 |
| Output | $12.00 | $6.00 |
| Cache read | $0.50 | $0.25 |
One route, standard clients.
Authenticate with a MiniRouter bearer key. Caller-provided upstream credentials are not used on this route.
Read model setup docs →https://api.minirouter.sh/v1alibaba/qwen3.8-max-prime/v1/chat/completions/v1/messages/v1/responsesmax_tokens, temperature, stop, tools, tool_choice, reasoning, include_reasoning, response_format, structured_outputs
JSON Schema and JSON mode enforced on verified providers.
Configured provider
Priority follows the live catalog order.
Input: text, image · Output: text
| Provider | Priority | Input / 1M | Output / 1M | Parameters | Access |
|---|---|---|---|---|---|
vercel | 1 | $4.2 | $12.6 | max_tokens, temperature, stop, tools, tool_choice, reasoning, include_reasoning, response_format, structured_outputs | MiniRouter balance |