Build
Tools and plugins
Use built-in tools, connect your own functions, or adjust prompts and responses with plugins.
Choose who runs the tool
- Server tools — MiniRouter gets the current time, searches the web or reads a public web page, then returns the model's answer in one request.
- Your own functions — your app runs code, such as a database query, and sends the result back to the model.
- Plugins — MiniRouter shortens a long prompt or repairs JSON. The model does not call a plugin.
The examples use Chat Completions (POST /v1/chat/completions).
MiniRouter's server tools and plugins are available on this endpoint only.
Tools require a model that supports function calling; check tools in its
supported_parameters in the catalog.
Try a server tool
Create an API key, save its recovery link, and add credits. Replace the placeholder below and run it in your terminal before running the examples on this page.
export MINIROUTER_KEY='paste-your-api-key-here'Examples use paid models. For free requests, follow the Auto Free guide.
Run this in the same terminal. It asks Qwen 3.5 Flash to get the current year
with minirouter:datetime.
curl https://api.minirouter.sh/v1/chat/completions \
-H "Authorization: Bearer $MINIROUTER_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "alibaba/qwen3.5-flash",
"messages": [
{
"role": "user",
"content": "Use the datetime tool to find the current year, then reply with only that four-digit year."
}
],
"max_tokens": 128,
"reasoning": {
"effort": "none"
},
"tools": [
{
"type": "minirouter:datetime"
}
],
"max_tool_calls": 2
}'Read the year in choices[0].message.content. You do not need to write a
datetime function or send a tool result yourself.
How the tool loop works
- You send a prompt and a tool. The
toolsarray tells the model what it can use. - The model requests a tool call. MiniRouter runs the built-in tool and sends its result back to the model.
- The model answers. MiniRouter returns that answer in the same HTTP response. It can repeat step 2 within your
max_tool_callslimit.
Listing a tool makes it available; the model may answer without using it.
Ask explicitly when the task needs a tool, as in the example above. In a
non-streaming response, x-minirouter-server-tool-rounds lists the model
request IDs. Two or more rounds show that the server tool loop continued.
You can mix built-in tools with your own function definitions. When the model requests one of your functions, MiniRouter returns that call for your app to handle using the function-calling flow.
Available server tools
| Tool | Use it for | Options |
|---|---|---|
minirouter:datetime | Current date and time | parameters.timezone: an IANA name such as Europe/London. Default: UTC. |
minirouter:web_fetch | Readable text from a public HTTP or HTTPS URL | Limit page size, domains and fetch count with parameters. Example below. |
minirouter:web_search | Current results from the web, $0.01 per search | Filter sources, freshness and language; limit search context. Search options. |
To set the timezone, replace the example's tools field with:
{
"tools": [{
"type": "minirouter:datetime",
"parameters": { "timezone": "Europe/London" }
}]
}Read a web page. This request allows one fetch from example.com:
curl https://api.minirouter.sh/v1/chat/completions \
-H "Authorization: Bearer $MINIROUTER_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "alibaba/qwen3.5-flash",
"messages": [
{
"role": "user",
"content": "Use web_fetch to read https://example.com and summarize the page in one sentence."
}
],
"max_tokens": 256,
"reasoning": {
"effort": "none"
},
"tools": [
{
"type": "minirouter:web_fetch",
"parameters": {
"allowed_domains": [
"example.com"
],
"max_uses": 1,
"max_content_tokens": 2000
}
}
],
"max_tool_calls": 2
}'Web fetch reads a URL; it does not search the web or sign in to a website. Private-network addresses are blocked. The page text is sent to the model as tool output, so treat it as untrusted input.
| Web fetch option | Meaning |
|---|---|
max_uses | Fetch attempts per API request, from 1 to 50. The overall max_tool_calls limit also applies. |
max_content_tokens | Approximate limit on returned page text, from 100 to 100,000. Default: 10,000. |
allowed_domains | Only allow these domains. Use hostnames, such as example.com, without a URL scheme or path. |
blocked_domains | Block these domains, even if they also match the allowlist. |
Search the web
Search Formula 1's site for results published in the past month:
curl https://api.minirouter.sh/v1/chat/completions \
-H "Authorization: Bearer $MINIROUTER_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "zai/glm-5.3-flash",
"messages": [
{
"role": "user",
"content": "Search the web for the latest Formula 1 Grand Prix result. Give the winner and source URL."
}
],
"tools": [
{
"type": "minirouter:web_search",
"parameters": {
"max_uses": 1,
"max_results": 3,
"search_domain_filter": [
"formula1.com"
],
"search_recency_filter": "month",
"country": "GB",
"search_language_filter": [
"en"
],
"max_tokens": 4096,
"max_tokens_per_page": 1024
}
}
]
}'Searches cost $0.01 each, plus model tokens. usage.server_tool_use reports
the billed count in web_search_requests, capped at max_uses.
The model can exceed this allowance; extra searches cost about $0.005 each.
Web search is powered by Perplexity. It currently requires a model served through
Vercel; other routes return 400.
Set these options inside the search tool's parameters. All are optional;
numeric values must be integers.
| Search option | Meaning |
|---|---|
max_uses | Billed allowance per API request: 1–10 searches at $0.01 each. Default: 3. The model can exceed it. |
max_results | Results per search, from 1 to 20. Omit to let the model choose. |
search_domain_filter | Up to 20 hostnames, without a scheme or path. Include with ["formula1.com"] or exclude with ["-reddit.com"]. Do not mix inclusion and exclusion. |
search_recency_filter | Publication recency: day, week, month or year. |
country | Regional results. Use an uppercase, two-letter ISO country code, such as GB or US. |
search_language_filter | Up to 10 lowercase, two-letter ISO language codes, such as ["en", "fr"]. |
max_tokens | Total search-result tokens per search, from 1 to 1,000,000. Default: 25,000. |
max_tokens_per_page | Search-result tokens per page, from 256 to 2,048. Default: 2,048. |
The tool's parameters.max_tokens bounds retrieved search context.
The request's top-level max_tokens limits the model's output per round.
Streaming, limits and cost
- Limit datetime and web-fetch calls:
max_tool_callsaccepts 1–30, defaults to 5, and requires at least one server tool. After reaching the limit, MiniRouter asks the model to finish without further tool calls. This does not cap web searches. - Allow for multiple model rounds: the model reads the prompt, then reads the tool results. Each round is billed separately;
max_tokenslimits each round, not the whole loop. - Read the total: non-streaming
usageandx-minirouter-cost-usdinclude all model rounds. For streaming totals, sendstream_options: {"include_usage":true}and read the final usage chunk. - Keep the connection open: use
stream: trueandcurl -Nfor longer loops. MiniRouter sends keep-alives while tools run, then sends the completed answer as SSE events. Intermediate tool rounds are not streamed token by token. Handle error events as well as the final[DONE].
For example, add these fields to the datetime request:
{
"stream": true,
"stream_options": { "include_usage": true }
}Run your own functions
Use a function tool when the work belongs in your app, such as querying an order or calling a weather API. A function definition describes the input; it does not upload or run your code.
curl https://api.minirouter.sh/v1/chat/completions \
-H "Authorization: Bearer $MINIROUTER_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "alibaba/qwen3.5-flash",
"reasoning": {
"effort": "none"
},
"messages": [
{
"role": "user",
"content": "What is the weather in London?"
}
],
"max_tokens": 256,
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city.",
"parameters": {
"type": "object",
"properties": {
"city": {
"type": "string"
}
},
"required": [
"city"
],
"additionalProperties": false
}
}
}
],
"tool_choice": "auto"
}'Then complete the conversation in your application:
- Read
choices[0].message.tool_calls. Withtool_choice: "auto", the model may answer directly instead. - Parse each function's
argumentsJSON, validate it, check permissions, and run your function. - Append the complete assistant message, including its
tool_calls, to your originalmessagesarray. - Append a
role: "tool"message for each result. Match itstool_call_idto the actual call ID and encodecontentas a string. - Send the updated conversation to the same endpoint and model to get the answer. Repeat if the model requests another call.
An illustrative tool result message (replace the ID and content with your result):
{
"role": "tool",
"tool_call_id": "call_id_from_response",
"content": "18 C, cloudy"
}On models that support tool choice, tool_choice: "required" forces a call
and "none" disables calls. parallel_tool_calls controls parallel
function calls where supported. Your app controls its own function-call limit;
max_tool_calls is for MiniRouter's built-in tools.
Plugins
Plugins are optional processing steps, enabled through plugins rather than
tools. They do not add functions for the model to call. You can include both
plugin IDs in one request; omit a plugin or set "enabled": false to disable it.
Fit a long prompt
Add this field to a Chat Completions request with an explicit model ID:
{
"plugins": [{ "id": "context-compression" }]
}If the prompt is too long for the model's context window after allowing for output, MiniRouter removes conversation turns from the middle, then shortens long messages if needed. It preserves the system-message prefix and the first and latest turns from removal, but message content may still be shortened.
This can discard details; it does not create a summary. Keep essential facts
in a short system message or the latest prompt, and keep your own complete
history. A prompt that already fits is unchanged. Use a named model, not
minirouter/auto, so MiniRouter can look up its context window.
Repair JSON output
Add these fields to a non-streaming request and ask for JSON in your prompt:
{
"response_format": { "type": "json_object" },
"plugins": [{ "id": "response-healing" }],
"stream": false
}The plugin can remove code fences and surrounding text, and close unfinished
brackets in choices[0].message.content. It also accepts a json_schema
response format; see structured output
for model support.
Valid JSON and output cut off by the token limit are returned unchanged.
Repair is best-effort, does not guarantee schema compliance, and does not run
with streaming. Check finish_reason, parse the JSON and validate it in your app.