Open weights, one field away, cheaper past 2K tokens.

Same key, same base URL. Darkbloom serves quantized open-weight models on Apple Silicon at a fraction of the default rates, with a floor on every request.

Base URLhttps://api.minirouter.sh/v1
Get a key — no account

Models on Darkbloom

6

4 only here

Input rate, best case

3.6× lower

Google Gemma 4 26B A4B

Floor per request

$0.000105

whenever tokens cost less

Precision

2-bit / 4-bit / 8-bit

named per model

How it works

One line in the request. One header back.

the request · chat completions or the Responses API
curl https://api.minirouter.sh/v1/chat/completions \  -H "Authorization: Bearer $MINIROUTER_KEY" \  -H "Content-Type: application/json" \  -d '{    "model": "openai/gpt-oss-20b",    "messages": [{"role": "user", "content": "ping"}],added line:     "providerOptions": {"gateway": {"only": ["darkbloom"]}}  }'
the response
HTTP/2 200added line: x-minirouter-provider: openai/gpt-oss-20b@darkbloomx-minirouter-cost-usd: 0.000105x-minirouter-request-id: req_…

Routing

Pinned means pinned. Unpinned means unchanged.

Where a request goesA request with providerOptions.gateway.only set to darkbloom is served on Darkbloom, or refused with 503 and nothing charged. A request without the field goes to the default route when the model has one, and to Darkbloom when only Darkbloom serves the model.servedunavailabledefault route existsDarkbloom onlyYour requestPOST /chat/completionsField setonly: ["darkbloom"]No fieldrequest unchangedDarkbloom@darkbloomRefused503, nothing chargedDefault route@vercelDarkbloom@darkbloom

Rates

Per million tokens, side by side.

After markup, from the live catalog. The default route is what an unpinned request pays.

6 models with a Darkbloom route, current catalog rates
ModelInput, $ per 1M tokensOutput, $ per 1M tokens
Qwen3 VL 30B A3B Instruct4-bit on Darkbloom
$0.095
—
Darkbloom only
$0.42
—
Darkbloom only
Qwen 3.5 9B4-bit on Darkbloom
$0.084
—
Darkbloom only
$0.137
—
Darkbloom only
Qwen 3.6 35B A3B4-bit on Darkbloom
$0.053
—
Darkbloom only
$0.735
—
Darkbloom only
Google Gemma 4 26B A4B4-bit on Darkbloom
$0.044
$0.158
3.6× lower on Darkbloom
$0.231
$0.63
2.7× lower on Darkbloom
GPT OSS 20B8-bit on Darkbloom
$0.021
$0.032
1.5× lower on Darkbloom
$0.105
$0.147
1.4× lower on Darkbloom
Bonsai 2 27B2-bit on Darkbloom
$0.079
—
Darkbloom only
$0.525
—
Darkbloom only
DarkbloomDefault routeOne scale across both columns.

The floor

Cheaper past 2K input tokens.

GPT OSS 20B, reply held at 250 tokens. Below the crossing the default route wins; above it Darkbloom does, and the gap widens.

Cost per request on GPT OSS 20B, 250 output tokens, by prompt sizeCost per request on GPT OSS 20B as the prompt grows from 0 to 8K input tokens with 250 output tokens. Darkbloom is flat at its $0.000105 floor until about 2K input tokens, where the default route costs the same; beyond that Darkbloom is cheaper.$0.000078$0.000156$0.000234$0.00031202K4K6K8Kinput tokens per request · output fixed at 250floor $0.000105cheaper past 2K tokensdefault routeDarkbloom
Cost per request on GPT OSS 20B, 250 output tokens, by prompt size, as a table
Input tokensDarkbloomDefault routeCheaper
0$0.000105$0.000037Default
250$0.000105$0.000045Default
500$0.000105$0.000053Default
1K$0.000105$0.000068Default
2K$0.000105$0.000100Default
4K$0.000110$0.000163Darkbloom
8K$0.000194$0.000289Darkbloom

Four workloads on GPT OSS 20B

Health check

A ping.

+7592%

20 in · 5 out

Darkbloom
$0.000105floor
Default
$0.000001

Chat turn

Short history, short reply.

+156%

600 in · 150 out

Darkbloom
$0.000105floor
Default
$0.000041

Agent turn

System prompt, tools, history.

−32%

4K in · 250 out

Darkbloom
$0.000110
Default
$0.000163

Document pass

A long file in, a summary out.

−33%

30K in · 500 out

Darkbloom
$0.000682
Default
$0.0010

Pin the real work; leave health checks and empty pings on the default route.

Next

Where to go from here.