darkbloom

Darkbloom

Open-weight models on Apple Silicon. Quantized serving, with a floor price on every request.

6 models

Serving
Apple Silicon, alpha

Open weights, quantized. Output can differ from full precision.

Precision
2-bit / 4-bit / 8-bit

Named per model in the comparison below.

Floor per request
$0.000105

Charged whenever the token cost is below it.

Select it
One request field

providerOptions.gateway.only; rules and arithmetic in the guide.

Your prompt
Sent to Darkbloom

30-day trace here; their retention in the privacy policy.

Rates against the default route

Per million tokens after markup, from the live catalog. The default route is what an unpinned request is billed at. The floor makes the comparison depend on request size; the guide draws where it stops mattering.

6 models with a Darkbloom route, current catalog rates
ModelInput, $ per 1M tokensOutput, $ per 1M tokens
Qwen3 VL 30B A3B Instruct4-bit on Darkbloom
$0.095
—
Darkbloom only
$0.42
—
Darkbloom only
Qwen 3.5 9B4-bit on Darkbloom
$0.084
—
Darkbloom only
$0.137
—
Darkbloom only
Qwen 3.6 35B A3B4-bit on Darkbloom
$0.053
—
Darkbloom only
$0.735
—
Darkbloom only
Google Gemma 4 26B A4B4-bit on Darkbloom
$0.044
$0.158
3.6× lower on Darkbloom
$0.231
$0.63
2.7× lower on Darkbloom
GPT OSS 20B8-bit on Darkbloom
$0.021
$0.032
1.5× lower on Darkbloom
$0.105
$0.147
1.4× lower on Darkbloom
Bonsai 2 27B2-bit on Darkbloom
$0.079
—
Darkbloom only
$0.525
—
Darkbloom only
DarkbloomDefault routeOne scale across both columns.

Available models

Author
Input
Providers
Limits

Context

Input price

Only

Sort

6 models

  • Bonsai 2 27B
    Context
    262K
    Input / 1M
    $0.0788
    Output / 1M
    $0.525
    Max output
    33K
    Accepts
    • Text
  • Qwen 3.6 35B A3B
    Context
    262K
    Input / 1M
    $0.0525
    Output / 1M
    $0.735
    Max output
    16K
    Accepts
    • Text
    • Image
  • Google Gemma 4 26B A4B
    Context
    262K
    Input / 1M
    $0.158
    Output / 1M
    $0.63
    Max output
    131K
    Accepts
    • Text
    • Image
    • PDF
  • Qwen 3.5 9B
    Context
    262K
    Input / 1M
    $0.084
    Output / 1M
    $0.137
    Max output
    16K
    Accepts
    • Text
    • Image
  • Qwen3 VL 30B A3B Instruct
    Context
    131K
    Input / 1M
    $0.0945
    Output / 1M
    $0.42
    Max output
    33K
    Accepts
    • Text
    • Image
  • GPT OSS 20B
    Context
    131K
    Input / 1M
    $0.0315
    Output / 1M
    $0.147
    Max output
    8K
    Accepts
    • Text