PrismML

Bonsai 2 27B

PrismML's Qwen3.8-derived 27B model, served by Darkbloom in native MLX 2-bit format. Start with text chat and evaluate it on your own coding and reasoning tasks.

Ternary weightsApache 2.0text inputtext output

Run this model

prism-ml/ternary-bonsai-2-27b

Use an endpoint supported by this model.

Get a key for this setup →
const response = await fetch('https://api.minirouter.sh/v1/chat/completions', {  method: 'POST',  headers: {    Authorization: `Bearer ${process.env.MINIROUTER_KEY}`,    'Content-Type': 'application/json',  },  body: JSON.stringify({    model: 'prism-ml/ternary-bonsai-2-27b',    messages: [{ role: 'user', content: 'Why is the sky blue?' }],    max_tokens: 1024,  }),}) const data = await response.json()console.log(data.choices[0].message.content)
Read docs →

Use prism-ml/ternary-bonsai-2-27b with Chat Completions. Image input and structured-output modes are not supported on this route. Reasoning can consume the output allowance before visible text appears; use a generous max_tokens limit.

Text input and a per-request minimum: Darkbloom serves ternary weights in MLX 2-bit format. This route accepts text input only. MiniRouter charges a minimum of $0.000105 per request on this model route; reasoning tokens count as output. Source ↗

Model overview

Input / 1M
$0.07875
All-in MiniRouter rate
Output / 1M
$0.525
All-in MiniRouter rate
Context
262K
262,144 tokens
Max output
33K
tokens
Foundation
27B dense
Derived from Qwen3.8 27B
Serving format
MLX 2-bit
Ternary language weights
Announced
17 Sept 2026
publisher release

Bonsai's second generation

A smaller representation of a 27B model.

Bonsai 2 uses three language-weight values with scaling. Darkbloom serves those weights in an MLX 2-bit container; this is a separate PrismML model ID, not an alias for ordinary Qwen3.8.

Read PrismML's announcement
  1. 01

    Keep format claims specific

    PrismML also distributes GGUF files. Their download sizes describe a different package from the MLX weights Darkbloom serves.

  2. 02

    Treat publisher scores as a starting point

    The announcement reports 98.2% aggregate benchmark retention against full precision. That result does not guarantee the same accuracy on your prompts or the same speed through a hosted API.

Evaluate on real tasks

Try bounded coding and reasoning work.

Compare task success and total cost using the same prompts, tool definitions, and completion criteria as your current model.

  1. 01

    Code explanations and small fixes

    Begin with tasks you can check: explain a function, propose a patch, or identify an error. Run the resulting code and tests before adopting an answer.

  2. 02

    Text documents and tool workflows

    The catalog above shows the route's limits. Build up context gradually; a maximum context window is not a latency or capacity guarantee. Validate tool arguments in your application.

  3. 03

    Choose another route for images

    The source model has vision, but this MiniRouter route accepts text only. Use a model that advertises image input for screenshots and visual documents.

Darkbloom billing

Include the minimum when comparing short requests.

The price panel reads MiniRouter's current catalog. Darkbloom bills input and output with a per-request minimum; reasoning tokens count toward output usage.

Understand Darkbloom pricing
  1. 01

    Tiny requests still have a floor

    The current minimum charge on this model route is $0.000105 per request. A token-only estimate below that floor understates the charge.

  2. 02

    Budget for reasoning and answers together

    Longer reasoning can change the total cost even when the final answer is short. Compare completed tasks, including retries, instead of input price alone.

  3. 03

    Keep the same MiniRouter key

    Select the exact model ID in the setup example. To require this provider explicitly, set providerOptions.gateway.only to ["darkbloom"] in a Chat Completions request.

Release and serving facts

Separate the model from the service.

The publisher's local hardware tests describe the weights and runtime. They do not measure MiniRouter latency or Darkbloom's network capacity.

Bonsai 2 release notes and sources
  1. 01

    Model release

    PrismML announced Bonsai 2 on September 17, 2026. Our release article links the primary announcement and the provider pricing feed.

  2. 02

    Hosted availability

    Model discovery and the route table reflect MiniRouter's published catalog. Supported inputs and limits can differ from a locally hosted copy of the model.

One route, standard clients.

Authenticate with a MiniRouter bearer key. Caller-provided upstream credentials are not used on this route.

Read model setup docs →
Base URLhttps://api.minirouter.sh/v1
Model IDprism-ml/ternary-bonsai-2-27b
Endpoints
/v1/chat/completions/v1/messages/v1/responses
Parameters

max_tokens, temperature, top_p, frequency_penalty, presence_penalty, stop, seed, tools, tool_choice

Structured output

Not enforced; a requested format is dropped and the answer is plain text.

Configured provider

Priority follows the live catalog order.

Input: text · Output: text

ProviderPriorityInput / 1MOutput / 1MParametersAccess
darkbloom
102$0.07875$0.525max_tokens, temperature, top_p, frequency_penalty, presence_penalty, stop, seed, tools, tool_choiceMiniRouter balance