Build

Tools and plugins

Use built-in tools, connect your own functions, or adjust prompts and responses with plugins.

Choose who runs the tool

  • Server tools — MiniRouter gets the current time, searches the web or reads a public web page, then returns the model's answer in one request.
  • Your own functions — your app runs code, such as a database query, and sends the result back to the model.
  • Plugins — MiniRouter shortens a long prompt or repairs JSON. The model does not call a plugin.

The examples use Chat Completions (POST /v1/chat/completions). MiniRouter's server tools and plugins are available on this endpoint only. Tools require a model that supports function calling; check tools in its supported_parameters in the catalog.

Try a server tool

Create an API key, save its recovery link, and add credits. Replace the placeholder below and run it in your terminal before running the examples on this page.

Try a server tool
export MINIROUTER_KEY='paste-your-api-key-here'

Examples use paid models. For free requests, follow the Auto Free guide.

Run this in the same terminal. It asks Qwen 3.5 Flash to get the current year with minirouter:datetime.

Try a server tool
curl https://api.minirouter.sh/v1/chat/completions \
  -H "Authorization: Bearer $MINIROUTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "alibaba/qwen3.5-flash",
  "messages": [
    {
      "role": "user",
      "content": "Use the datetime tool to find the current year, then reply with only that four-digit year."
    }
  ],
  "max_tokens": 128,
  "reasoning": {
    "effort": "none"
  },
  "tools": [
    {
      "type": "minirouter:datetime"
    }
  ],
  "max_tool_calls": 2
}'

Read the year in choices[0].message.content. You do not need to write a datetime function or send a tool result yourself.

How the tool loop works

  1. You send a prompt and a tool. The tools array tells the model what it can use.
  2. The model requests a tool call. MiniRouter runs the built-in tool and sends its result back to the model.
  3. The model answers. MiniRouter returns that answer in the same HTTP response. It can repeat step 2 within your max_tool_calls limit.

Listing a tool makes it available; the model may answer without using it. Ask explicitly when the task needs a tool, as in the example above. In a non-streaming response, x-minirouter-server-tool-rounds lists the model request IDs. Two or more rounds show that the server tool loop continued.

You can mix built-in tools with your own function definitions. When the model requests one of your functions, MiniRouter returns that call for your app to handle using the function-calling flow.

Available server tools

ToolUse it forOptions
minirouter:datetimeCurrent date and timeparameters.timezone: an IANA name such as Europe/London. Default: UTC.
minirouter:web_fetchReadable text from a public HTTP or HTTPS URLLimit page size, domains and fetch count with parameters. Example below.
minirouter:web_searchCurrent results from the web, $0.01 per searchFilter sources, freshness and language; limit search context. Search options.

To set the timezone, replace the example's tools field with:

Available server tools
{
  "tools": [{
    "type": "minirouter:datetime",
    "parameters": { "timezone": "Europe/London" }
  }]
}

Read a web page. This request allows one fetch from example.com:

Available server tools
curl https://api.minirouter.sh/v1/chat/completions \
  -H "Authorization: Bearer $MINIROUTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "alibaba/qwen3.5-flash",
  "messages": [
    {
      "role": "user",
      "content": "Use web_fetch to read https://example.com and summarize the page in one sentence."
    }
  ],
  "max_tokens": 256,
  "reasoning": {
    "effort": "none"
  },
  "tools": [
    {
      "type": "minirouter:web_fetch",
      "parameters": {
        "allowed_domains": [
          "example.com"
        ],
        "max_uses": 1,
        "max_content_tokens": 2000
      }
    }
  ],
  "max_tool_calls": 2
}'

Web fetch reads a URL; it does not search the web or sign in to a website. Private-network addresses are blocked. The page text is sent to the model as tool output, so treat it as untrusted input.

Web fetch optionMeaning
max_usesFetch attempts per API request, from 1 to 50. The overall max_tool_calls limit also applies.
max_content_tokensApproximate limit on returned page text, from 100 to 100,000. Default: 10,000.
allowed_domainsOnly allow these domains. Use hostnames, such as example.com, without a URL scheme or path.
blocked_domainsBlock these domains, even if they also match the allowlist.

Search the web

Search Formula 1's site for results published in the past month:

Search the web
curl https://api.minirouter.sh/v1/chat/completions \
  -H "Authorization: Bearer $MINIROUTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "zai/glm-5.3-flash",
  "messages": [
    {
      "role": "user",
      "content": "Search the web for the latest Formula 1 Grand Prix result. Give the winner and source URL."
    }
  ],
  "tools": [
    {
      "type": "minirouter:web_search",
      "parameters": {
        "max_uses": 1,
        "max_results": 3,
        "search_domain_filter": [
          "formula1.com"
        ],
        "search_recency_filter": "month",
        "country": "GB",
        "search_language_filter": [
          "en"
        ],
        "max_tokens": 4096,
        "max_tokens_per_page": 1024
      }
    }
  ]
}'

Searches cost $0.01 each, plus model tokens. usage.server_tool_use reports the billed count in web_search_requests, capped at max_uses. The model can exceed this allowance; extra searches cost about $0.005 each.

Web search is powered by Perplexity. It currently requires a model served through Vercel; other routes return 400.

Set these options inside the search tool's parameters. All are optional; numeric values must be integers.

Search optionMeaning
max_usesBilled allowance per API request: 1–10 searches at $0.01 each. Default: 3. The model can exceed it.
max_resultsResults per search, from 1 to 20. Omit to let the model choose.
search_domain_filterUp to 20 hostnames, without a scheme or path. Include with ["formula1.com"] or exclude with ["-reddit.com"]. Do not mix inclusion and exclusion.
search_recency_filterPublication recency: day, week, month or year.
countryRegional results. Use an uppercase, two-letter ISO country code, such as GB or US.
search_language_filterUp to 10 lowercase, two-letter ISO language codes, such as ["en", "fr"].
max_tokensTotal search-result tokens per search, from 1 to 1,000,000. Default: 25,000.
max_tokens_per_pageSearch-result tokens per page, from 256 to 2,048. Default: 2,048.

The tool's parameters.max_tokens bounds retrieved search context. The request's top-level max_tokens limits the model's output per round.

Streaming, limits and cost

  • Limit datetime and web-fetch calls: max_tool_calls accepts 1–30, defaults to 5, and requires at least one server tool. After reaching the limit, MiniRouter asks the model to finish without further tool calls. This does not cap web searches.
  • Allow for multiple model rounds: the model reads the prompt, then reads the tool results. Each round is billed separately; max_tokens limits each round, not the whole loop.
  • Read the total: non-streaming usage and x-minirouter-cost-usd include all model rounds. For streaming totals, send stream_options: {"include_usage":true} and read the final usage chunk.
  • Keep the connection open: use stream: true and curl -N for longer loops. MiniRouter sends keep-alives while tools run, then sends the completed answer as SSE events. Intermediate tool rounds are not streamed token by token. Handle error events as well as the final [DONE].

For example, add these fields to the datetime request:

Streaming, limits and cost
{
  "stream": true,
  "stream_options": { "include_usage": true }
}

Run your own functions

Use a function tool when the work belongs in your app, such as querying an order or calling a weather API. A function definition describes the input; it does not upload or run your code.

Run your own functions
curl https://api.minirouter.sh/v1/chat/completions \
  -H "Authorization: Bearer $MINIROUTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "alibaba/qwen3.5-flash",
  "reasoning": {
    "effort": "none"
  },
  "messages": [
    {
      "role": "user",
      "content": "What is the weather in London?"
    }
  ],
  "max_tokens": 256,
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get the current weather for a city.",
        "parameters": {
          "type": "object",
          "properties": {
            "city": {
              "type": "string"
            }
          },
          "required": [
            "city"
          ],
          "additionalProperties": false
        }
      }
    }
  ],
  "tool_choice": "auto"
}'

Then complete the conversation in your application:

  1. Read choices[0].message.tool_calls. With tool_choice: "auto", the model may answer directly instead.
  2. Parse each function's arguments JSON, validate it, check permissions, and run your function.
  3. Append the complete assistant message, including its tool_calls, to your original messages array.
  4. Append a role: "tool" message for each result. Match its tool_call_id to the actual call ID and encode content as a string.
  5. Send the updated conversation to the same endpoint and model to get the answer. Repeat if the model requests another call.

An illustrative tool result message (replace the ID and content with your result):

Run your own functions
{
  "role": "tool",
  "tool_call_id": "call_id_from_response",
  "content": "18 C, cloudy"
}

On models that support tool choice, tool_choice: "required" forces a call and "none" disables calls. parallel_tool_calls controls parallel function calls where supported. Your app controls its own function-call limit; max_tool_calls is for MiniRouter's built-in tools.

Plugins

Plugins are optional processing steps, enabled through plugins rather than tools. They do not add functions for the model to call. You can include both plugin IDs in one request; omit a plugin or set "enabled": false to disable it.

Fit a long prompt

Add this field to a Chat Completions request with an explicit model ID:

Fit a long prompt
{
  "plugins": [{ "id": "context-compression" }]
}

If the prompt is too long for the model's context window after allowing for output, MiniRouter removes conversation turns from the middle, then shortens long messages if needed. It preserves the system-message prefix and the first and latest turns from removal, but message content may still be shortened.

This can discard details; it does not create a summary. Keep essential facts in a short system message or the latest prompt, and keep your own complete history. A prompt that already fits is unchanged. Use a named model, not minirouter/auto, so MiniRouter can look up its context window.

Repair JSON output

Add these fields to a non-streaming request and ask for JSON in your prompt:

Repair JSON output
{
  "response_format": { "type": "json_object" },
  "plugins": [{ "id": "response-healing" }],
  "stream": false
}

The plugin can remove code fences and surrounding text, and close unfinished brackets in choices[0].message.content. It also accepts a json_schema response format; see structured output for model support.

Valid JSON and output cut off by the token limit are returned unchanged. Repair is best-effort, does not guarantee schema compliance, and does not run with streaming. Check finish_reason, parse the JSON and validate it in your app.

Esc