Operate

Guardrails

Set a budget, restrict model access, and filter request content with a policy attached to your API keys.

1. Set a budget

Give each key its own daily, weekly, or monthly spend limit.

2. Choose access

Control which models and providers a key can use.

3. Filter input

Flag, redact, or block sensitive text before it reaches a provider.

Input only. Content filters scan supported text requests, never model output. Embeddings, rerank, and image requests still follow budgets and access lists, but are not content-scanned. Regex filters are a limited check, not a complete defense against prompt injection.

Create it. Test it. Attach it.

Create a guardrail and test its filters, then select it on your API key. A key can have one guardrail; an account can have up to 20.

Assignment requires a signed-in browser session. An empty guardrail adds no restrictions. Deleting one removes it from every assigned key.

Assigning a guardrail to a key

Assign through the dashboard, or use the key API from a signed-in session.

Pick it in the Guardrail select on /dashboard/keys, or send guardrailId (null to clear) on POST /keys or PATCH /keys/:id.

Assign from a signed-in session
PATCH https://api.minirouter.sh/keys/42
{ "guardrailId": "9b1c2d3e-4f50-4a6b-8c7d-0e1f2a3b4c5d" }

Only a signed-in browser session may assign one. An inference key sending guardrailId is refused with 403, because swapping a guardrail could widen a model list or drop a filter. Key views return guardrailId and guardrailName. Deleting a guardrail unassigns every key that used it.

Budget

Budgets apply per key, include in-progress requests, and reset in UTC.

limitNano (nano-USD, decimal string) and resetInterval (daily, weekly, monthly) are set together or both null.

The budget is counted per key: one guardrail on three keys gives each key its own budget, never a pooled one. Windows are UTC: daily resets at 00:00, weekly on Monday 00:00, monthly on the 1st. In-progress requests count against the budget.

A request that would cross the budget is refused before the first byte with 403 spend_limit_exceeded. The message names the limit, the scope guardrail, the instant it resets (resetsAt) and where to change it. There is no Retry-After; retrying before the reset repeats the same error.

A budget crossed mid-stream ends the stream with the existing usage_limit_exceeded event, because bytes have already been sent.

How policies combine

Allowlists intersect. Ignored lists subtract. Every applicable policy still applies.

Allowlists intersect: the account list, the key’s allowedModelsand the guardrail’s allowedModels all apply, and a model must be in every non-empty one. Ignored lists subtract. An empty list on the guardrail adds no restriction. The same holds for providers.

A model that exists but falls outside the result returns 403 model_not_allowed, naming the model and the nearest allowed one when there is a close match. minirouter/autoskips the guardrail’s ignored models when it picks.

A model left with no provider after the provider lists are applied is refused the way an over-narrow key provider allowlist is refused today.

Sensitive info presets

Flag, redact, or block email addresses, secrets, and other sensitive text.

Each built-in detector has an action: flag records the hit, redact replaces each match with its placeholder before the request is measured, held or sent, and block refuses the request. A detector you do not list is off. Presets scan every message role regardless of the injection scan scope.

SlugDetectsPlaceholderExample
emailEmail address[EMAIL]user@example.com → [EMAIL]
phonePhone number[PHONE]914-309-4996 → [PHONE]
ssnSocial Security number[SSN]123-45-6789 → [SSN]
credit-cardCredit card number[CREDIT_CARD]4111 1111 1111 1111 → [CREDIT_CARD]
ip-addressIP address[IP_ADDRESS]203.0.113.7 → [IP_ADDRESS]
secretsSecret[SECRET:<format-id>]AKIAIOSFODNN7EXAMPLE → [SECRET:aws-access-key-id]
regex-prompt-injectionPrompt injection[PROMPT_INJECTION]Ignore all previous instructions → [PROMPT_INJECTION]

phone matches any run of ten digits with or without separators, so order numbers of that shape are redacted too. ip-address is IPv4 only. credit-card keeps only digit strings that pass the Luhn check.

The secrets preset is one regex per format. A hit names the format id in events and in the placeholder.

Format idDetectsPlaceholder
aws-access-key-idAWS access key ID[SECRET:aws-access-key-id]
github-tokenGitHub token[SECRET:github-token]
github-fine-grained-patGitHub fine-grained token[SECRET:github-fine-grained-pat]
gitlab-personal-access-tokenGitLab personal access token[SECRET:gitlab-personal-access-token]
openai-api-keyOpenAI API key[SECRET:openai-api-key]
anthropic-api-keyAnthropic API key[SECRET:anthropic-api-key]
openrouter-api-keyOpenRouter API key[SECRET:openrouter-api-key]
minirouter-api-keyMiniRouter API key[SECRET:minirouter-api-key]
google-api-keyGoogle API key[SECRET:google-api-key]
stripe-secret-keyStripe secret key[SECRET:stripe-secret-key]
slack-tokenSlack token[SECRET:slack-token]
slack-webhook-urlSlack webhook URL[SECRET:slack-webhook-url]
npm-access-tokennpm access token[SECRET:npm-access-token]
sendgrid-api-keySendGrid API key[SECRET:sendgrid-api-key]
huggingface-access-tokenHugging Face access token[SECRET:huggingface-access-token]
json-web-tokenJSON Web Token[SECRET:json-web-token]
digitalocean-tokenDigitalOcean token[SECRET:digitalocean-token]
telegram-bot-tokenTelegram bot token[SECRET:telegram-bot-token]
private-key-blockPrivate key block[SECRET:private-key-block]

Prompt injection

Built-in patterns, scan scope, and detection limits.

The regex-prompt-injection preset carries the 33 patterns below, ported from OpenRouter’s published set so the names in your events match the names in their documentation. Actions are flag, redact (placeholder [PROMPT_INJECTION]) and block.

injectionScanScope decides which text the injection patterns read. user_only (the default) reads user turns. all_messages also reads system prompts, assistant turns and tool arguments, including instructions a preset injected, because the scan runs after presets shape the body.

Spaced letters are collapsed first: three or more single letters in a row (i g n o r e) are read as one word and the patterns run again on the collapsed copy. A hit there counts under the same pattern name, and redaction replaces the original spaced span.

Three patterns are bounded so no repeat is unbounded on a 32 MB body: dan_jailbreak requires “do anything now” within 500 characters of DAN, identity_hijack caps the hijacked name at 64 word characters, and system_prefix_spoofing does not include blank lines before System:.

Not detected: scrambled words (typoglycemia), misspellings beyond the few in reveal_prompt and show_prompt, and text encoded as base64 or hex.

Direct Instruction Override
PatternDetects
ignore_previous_instructionsAttempts to discard prior instructions, optionally scoped to safety/system/etc.
disregard_instructionsVariants of "disregard your instructions/rules/guidelines/constraints/directives".
forget_instructionsAttempts to erase prior instructions/rules/guidelines/constraints/directives.
new_instructionsInjection marker introducing replacement instructions.
do_not_followTelling the model to disobey its system prompt.
supersede_instructions"Supersedes prior instructions" override.
void_instructionsClaims prior instructions are void/invalid/revoked/cancelled.
Developer / Admin Mode Activation
PatternDetects
developer_modeClaims the model is in developer mode.
enter_special_modeRequests to enter a special (developer/admin/debug/maintenance) mode.
activate_special_modeRequests to activate a special (developer/admin/debug/jailbreak) mode.
System Override
PatternDetects
system_overrideDirect system-override keyword.
override_instructionsAttempts to override instructions/rules/guidelines/constraints/directives.
Prompt Extraction
PatternDetects
reveal_promptAsks the model to reveal its (full/hidden/internal/secret/original/…) prompt.
show_promptAsks the model to show its prompt.
what_instructionsAsks what the model's instructions are.
repeat_instructionsAsks the model to repeat earlier text.
output_promptAsks for the original system prompt.
Role Manipulation
PatternDetects
remove_restrictionsClaims the model is no longer restricted.
act_unboundAsks the model to pretend it has no restrictions.
pretend_differentAsks the model to impersonate a different AI.
identity_hijackIdentity hijacking with explicitly malicious modifiers.
DAN-Style Jailbreaks
PatternDetects
dan_jailbreakThe classic DAN jailbreak (case-sensitive for "DAN").
jailbreak_modeReferences to jailbreak modes or prompts.
Safety Bypass
PatternDetects
bypass_safetyAttempts to bypass safety/security/content/ethical filters.
disable_safetyAttempts to disable safety/security/content measures.
ignore_safetyAttempts to ignore or disregard safety/security/ethical/content guidelines, rules, or restrictions.
Tag Injection & Role Spoofing
PatternDetects
system_tag_injectionInjecting `<system>`, `</system>`, or `<system/>` tags.
role_tag_injectionInjecting role-related XML tags (including self-closing).
role_delimiter_injectionInjecting role delimiters like `[system]:`.
bracketed_role_spoofingFake bracketed role labels (e.g. `[System]`, `[Assistant]`).
system_prefix_spoofingLines starting with `System:` to impersonate system messages (multiline).
Control Token Injection
PatternDetects
control_token_injectionChatML / Llama 3 / generic pipe-delimited control tokens.
deepseek_control_token_injectionDeepSeek fullwidth-pipe (`|`) control tokens.

Custom patterns

Add your own regular expressions and redaction labels.

contentFilters holds up to 20 entries of { pattern, action, label }. pattern is a JavaScript regular expression source of 1 to 2,000 characters, compiled with the g and u flags. Matching is case-sensitive: write [Aa]cme to match either case.

label is optional, up to 40 letters, digits, spaces, _ or -. The redaction placeholder is [label], or [REDACTED] without a label. Events list the label, or pattern n for an unlabeled entry.

A pattern is rejected with 400 when it uses lookahead, lookbehind or backreferences (\1, \k<name>); nests a quantifier inside a quantified group such as (a+)+; matches the empty string; or runs too slowly on plain text in the save-time probe. The message reads Pattern n: … and never repeats the pattern. Custom patterns scan every message role.

What is scanned

Supported endpoints, fields, and when redaction happens.

Only these strings are read and rewritten, on /chat/completions, /responses, /messages and /messages/count_tokens. Anything else at these paths, and every other field, is left alone.

RoutePathRole for scan scope
/chat/completionsmessages[i].content when it is a stringmessages[i].role
messages[i].content[j].text when content[j].type is textmessages[i].role
messages[i].tool_calls[j].function.arguments when it is a stringassistant
/responsesinput when it is a stringuser
input[i].content when it is a stringinput[i].role, else user
input[i].content[j].text when the type is input_text or output_textinput[i].role, else user
input[i].arguments when input[i].type is function_callassistant
instructions when it is a stringsystem
/messages, /messages/count_tokenssystem when it is a string; system[i].text when system[i].type is textsystem
messages[i].content when it is a stringmessages[i].role
messages[i].content[j].text when the type is textmessages[i].role
messages[i].content[j].content when the type is tool_result and it is a stringuser

The scan runs before token measurement, the balance hold, the request trace and the upstream send, so all of them see the redacted text.

Input only: model output is never scanned. Embeddings, rerank and image requests are not content-scanned; they still get the merged model and provider lists and the budget.

When a request is blocked

A content block returns 403 before a balance hold or provider request. No charge.

Block filters run first. Any hit refuses the request with 403 guardrail_blocked before anything is measured, held or sent, and nothing is charged. The message names the detector or pattern, never the text it matched, and is capped at 200 characters.

403 on /chat/completions and /responses
{
  "error": {
    "message": "A guardrail on this API key blocked the request: prompt injection patterns detected (ignore_previous_instructions). Nothing was sent to a provider and nothing was charged. Manage guardrails at https://minirouter.sh/dashboard/guardrails Docs: https://minirouter.sh/docs/guardrails",
    "type": "invalid_request_error",
    "code": "guardrail_blocked",
    "request_id": "req_01J8Z6Q0GUARDRAIL"
  }
}
403 on /messages and /messages/count_tokens
{
  "type": "error",
  "error": {
    "type": "permission_error",
    "message": "A guardrail on this API key blocked the request: prompt injection patterns detected (ignore_previous_instructions). Nothing was sent to a provider and nothing was charged. Manage guardrails at https://minirouter.sh/dashboard/guardrails Docs: https://minirouter.sh/docs/guardrails"
  },
  "request_id": "req_01J8Z6Q0GUARDRAIL"
}

The hit appears in the hits list on /dashboard/guardrails and on GET /guardrails/events. Hits are kept for 30 days.

Headers

See whether a request was flagged, redacted, or skipped.

x-minirouter-guardrail is set when a guardrail acted without blocking. A block is a 403, never a header; a request nothing matched carries no header unless the scan failed (skipped).

ValueMeaning
redactedText was replaced before the request was measured, held against your balance or sent upstream.
flaggedA flag filter matched. The request went through unchanged and the hit was recorded.
skippedThe scanner failed on this request. It was served unscanned and the failure was logged.

API

Endpoints, configuration fields, examples, and event history.

Base URL https://api.minirouter.sh, the same as /keys. Reads accept an API key or a signed-in session; writes and the test endpoint need a signed-in session. Writes are limited to 30 per minute per account.

EndpointPrincipalBehaviour
GET /guardrailskey or sessionEvery guardrail on the account plus limits { maxGuardrails: 20, maxCustomPatterns: 20 }.
POST /guardrailssessionBody: name, description and any config field; omitted fields start empty. 201 with the guardrail. 409 for a duplicate name or a 21st guardrail.
GET /guardrails/:idkey or sessionOne guardrail. 404 when the id is not on this account.
PATCH /guardrails/:idsessionBody: expectedRevision plus the fields to change. Arrays replace, never merge. 409 policy_revision_conflict when the revision is stale.
DELETE /guardrails/:idsessionHard delete. Keys that used it are unassigned. Returns { deleted: true, unassignedKeys }.
POST /guardrails/testsessionBody: text (up to 20,000 characters) plus exactly one of guardrailId or config. Scans text as a user turn and returns { action, redacted, hits }. Nothing is stored or logged.
GET /guardrails/eventskey or sessionHits, newest first. Query: limit (1 to 200, default 50), cursor (the nextCursor of the previous page), guardrailId.
Create a guardrail
POST https://api.minirouter.sh/guardrails
{
  "name": "Support bot",
  "limitNano": "25000000000",
  "resetInterval": "weekly",
  "contentFilterBuiltins": [
    { "slug": "email", "action": "redact" },
    { "slug": "secrets", "action": "block" },
    { "slug": "regex-prompt-injection", "action": "block" }
  ],
  "contentFilters": [
    { "pattern": "PROJ-\\d{3,6}", "action": "redact", "label": "ticket" }
  ],
  "injectionScanScope": "user_only"
}
FieldMeaning
nameUnique on the account, up to 80 characters.
descriptionOptional, up to 400 characters.
limitNanoBudget in nano-USD as a decimal string, or null. Sent together with resetInterval.
resetIntervaldaily, weekly or monthly, or null.
allowedModelsCatalog ids or author/* families. Empty means every model. Up to 317 entries.
ignoredModelsSame grammar. These models are refused even when allowed elsewhere.
allowedProvidersProvider ids. Empty means every provider.
ignoredProvidersProvider ids to remove.
contentFilterBuiltinsArray of { slug, action }. One entry per detector; a detector not listed is off.
contentFiltersArray of { pattern, action, label }. Up to 20 patterns.
injectionScanScopeuser_only (default) or all_messages.

Every guardrail is returned with id, revision (send it back as expectedRevision), keyCount, createdAt and updatedAt. Unknown fields are a 400 naming the field.

Try text against an unsaved config
POST https://api.minirouter.sh/guardrails/test
{ "config": { "contentFilterBuiltins": [ { "slug": "email", "action": "redact" } ] },
  "text": "Reach me at user@example.com" }

{ "action": "redacted",
  "redacted": "Reach me at [EMAIL]",
  "hits": [ { "detector": "email", "action": "redact", "patterns": [], "matchCount": 1 } ] }

Events carry id, guardrailId, guardrailName (null once deleted), apiKeyId, keyPrefix, requestId, route, detector (a preset slug or custom), action, patterns (pattern names, format ids or labels) and matchCount. Pages return nextCursor; pass it as cursor for the next page. null means the end.

Esc