---
title: "Guardrails"
description: "Per-key budgets, model and provider lists, sensitive-info redaction, prompt-injection detection and custom patterns."
canonical_url: "https://minirouter.sh/docs/guardrails"
markdown_url: "https://minirouter.sh/docs/guardrails.md"
last_updated: "2026-09-26"
---

# Guardrails

A named policy on an API key: a spend budget, model and provider lists, and filters that redact or block sensitive text and prompt injection before a request leaves the edge.

## What a guardrail is

A guardrail is a named policy you assign to an API key. It can carry a spend
budget, allowed and ignored model lists, allowed and ignored provider lists,
built-in content filters and custom regex filters. Create one at
https://minirouter.sh/dashboard/guardrails or with the API below. Each key
takes at most one guardrail; up to 20 per account. A key without a guardrail
behaves exactly as before, and a guardrail with nothing configured changes
nothing.

## Assigning a guardrail to a key

Pick it in the Guardrail select on https://minirouter.sh/dashboard/keys, or
send guardrailId (null to clear) on POST /keys or PATCH /keys/:id. Only a
signed-in browser session may assign one: an inference key sending
guardrailId is refused with 403, because swapping a guardrail could widen a
model list or drop a filter. Key views return guardrailId and guardrailName.
Deleting a guardrail unassigns every key that used it.

## Budget

limitNano (nano-USD, decimal string) and resetInterval (daily, weekly,
monthly) are set together or both null. The budget is counted per key: one
guardrail on three keys gives each key its own budget, never a pooled one.
Windows are UTC: daily resets at 00:00, weekly on Monday 00:00, monthly on
the 1st. In-progress requests count against the budget.

A request that would cross the budget is refused before the first byte with
403 spend_limit_exceeded. The message names the limit (for example "weekly
spend limit"), the scope "guardrail", the instant it resets (resetsAt) and
https://minirouter.sh/dashboard/guardrails. There is no Retry-After;
retrying before the reset repeats the same error. A budget crossed
mid-stream ends the stream with the existing usage_limit_exceeded event,
because bytes have already been sent.
Error reference: https://minirouter.sh/docs/errors#spend_limit_exceeded

## How the key's own policy and the guardrail combine

Allowlists intersect: the account list, the key's allowedModels and the
guardrail's allowedModels are all applied, and a model must be in every
non-empty one. Ignored lists subtract. An empty list on the guardrail adds no
restriction. The same holds for providers.

A model that exists but falls outside the result returns 403
model_not_allowed, naming the model and the nearest allowed one when there
is a close match. minirouter/auto skips the guardrail's ignored models when
it picks. A model left with no provider after the provider lists are applied
is refused the way an over-narrow key provider allowlist is refused today.
Error reference: https://minirouter.sh/docs/errors#model_not_allowed

## Sensitive info presets

Each built-in detector has an action: flag (record only), redact (replace
each match with its placeholder before the request is measured, held or
sent) or block (refuse the request). A detector you do not list is off.
Presets scan every message role regardless of the injection scan scope.

| Slug | Detects | Placeholder | Example |
|---|---|---|---|
| email | Email address | [EMAIL] | user@example.com → [EMAIL] |
| phone | Phone number | [PHONE] | 914-309-4996 → [PHONE] |
| ssn | Social Security number | [SSN] | 123-45-6789 → [SSN] |
| credit-card | Credit card number | [CREDIT_CARD] | 4111 1111 1111 1111 → [CREDIT_CARD] |
| ip-address | IP address | [IP_ADDRESS] | 203.0.113.7 → [IP_ADDRESS] |
| secrets | Secret | [SECRET:<format-id>] | AKIAIOSFODNN7EXAMPLE → [SECRET:aws-access-key-id] |
| regex-prompt-injection | Prompt injection | [PROMPT_INJECTION] | Ignore all previous instructions → [PROMPT_INJECTION] |

phone matches any run of ten digits with or without separators, so order
numbers and identifiers of that shape are redacted too. ip-address is IPv4
only. credit-card keeps only digit strings that pass the Luhn check.

The secrets preset is one regex per format. A hit names the format id in
events and in [SECRET:<id>].

| Format id | Detects | Placeholder |
|---|---|---|
| aws-access-key-id | AWS access key ID | [SECRET:aws-access-key-id] |
| github-token | GitHub token | [SECRET:github-token] |
| github-fine-grained-pat | GitHub fine-grained token | [SECRET:github-fine-grained-pat] |
| gitlab-personal-access-token | GitLab personal access token | [SECRET:gitlab-personal-access-token] |
| openai-api-key | OpenAI API key | [SECRET:openai-api-key] |
| anthropic-api-key | Anthropic API key | [SECRET:anthropic-api-key] |
| openrouter-api-key | OpenRouter API key | [SECRET:openrouter-api-key] |
| minirouter-api-key | MiniRouter API key | [SECRET:minirouter-api-key] |
| google-api-key | Google API key | [SECRET:google-api-key] |
| stripe-secret-key | Stripe secret key | [SECRET:stripe-secret-key] |
| slack-token | Slack token | [SECRET:slack-token] |
| slack-webhook-url | Slack webhook URL | [SECRET:slack-webhook-url] |
| npm-access-token | npm access token | [SECRET:npm-access-token] |
| sendgrid-api-key | SendGrid API key | [SECRET:sendgrid-api-key] |
| huggingface-access-token | Hugging Face access token | [SECRET:huggingface-access-token] |
| json-web-token | JSON Web Token | [SECRET:json-web-token] |
| digitalocean-token | DigitalOcean token | [SECRET:digitalocean-token] |
| telegram-bot-token | Telegram bot token | [SECRET:telegram-bot-token] |
| private-key-block | Private key block | [SECRET:private-key-block] |

## Prompt injection

The regex-prompt-injection preset carries the 33 patterns below, ported
from OpenRouter's published set so the names in your events match the names
in their documentation. Actions are flag, redact (placeholder
[PROMPT_INJECTION]) and block.

injectionScanScope decides which text the injection patterns read: user_only
(the default) reads user turns; all_messages also reads system prompts,
assistant turns and tool arguments, including instructions a preset injected,
because the scan runs after presets shape the body.

Spaced letters are collapsed first: three or more single letters in a row
("i g n o r e") are read as one word and the patterns run again on the
collapsed copy. A hit there counts under the same pattern name, and redaction
replaces the original spaced span.

Three patterns are bounded so no repeat is unbounded on a 32 MB body:
dan_jailbreak requires "do anything now" within 500 characters of DAN,
identity_hijack caps the hijacked name at 64 word characters, and
system_prefix_spoofing does not include blank lines before "System:".

Not detected: scrambled words (typoglycemia), misspellings beyond the few in
reveal_prompt and show_prompt, and text encoded as base64 or hex.

### Direct Instruction Override

| Pattern | Detects |
|---|---|
| ignore_previous_instructions | Attempts to discard prior instructions, optionally scoped to safety/system/etc. |
| disregard_instructions | Variants of "disregard your instructions/rules/guidelines/constraints/directives". |
| forget_instructions | Attempts to erase prior instructions/rules/guidelines/constraints/directives. |
| new_instructions | Injection marker introducing replacement instructions. |
| do_not_follow | Telling the model to disobey its system prompt. |
| supersede_instructions | "Supersedes prior instructions" override. |
| void_instructions | Claims prior instructions are void/invalid/revoked/cancelled. |

### Developer / Admin Mode Activation

| Pattern | Detects |
|---|---|
| developer_mode | Claims the model is in developer mode. |
| enter_special_mode | Requests to enter a special (developer/admin/debug/maintenance) mode. |
| activate_special_mode | Requests to activate a special (developer/admin/debug/jailbreak) mode. |

### System Override

| Pattern | Detects |
|---|---|
| system_override | Direct system-override keyword. |
| override_instructions | Attempts to override instructions/rules/guidelines/constraints/directives. |

### Prompt Extraction

| Pattern | Detects |
|---|---|
| reveal_prompt | Asks the model to reveal its (full/hidden/internal/secret/original/…) prompt. |
| show_prompt | Asks the model to show its prompt. |
| what_instructions | Asks what the model's instructions are. |
| repeat_instructions | Asks the model to repeat earlier text. |
| output_prompt | Asks for the original system prompt. |

### Role Manipulation

| Pattern | Detects |
|---|---|
| remove_restrictions | Claims the model is no longer restricted. |
| act_unbound | Asks the model to pretend it has no restrictions. |
| pretend_different | Asks the model to impersonate a different AI. |
| identity_hijack | Identity hijacking with explicitly malicious modifiers. |

### DAN-Style Jailbreaks

| Pattern | Detects |
|---|---|
| dan_jailbreak | The classic DAN jailbreak (case-sensitive for "DAN"). |
| jailbreak_mode | References to jailbreak modes or prompts. |

### Safety Bypass

| Pattern | Detects |
|---|---|
| bypass_safety | Attempts to bypass safety/security/content/ethical filters. |
| disable_safety | Attempts to disable safety/security/content measures. |
| ignore_safety | Attempts to ignore or disregard safety/security/ethical/content guidelines, rules, or restrictions. |

### Tag Injection & Role Spoofing

| Pattern | Detects |
|---|---|
| system_tag_injection | Injecting `<system>`, `</system>`, or `<system/>` tags. |
| role_tag_injection | Injecting role-related XML tags (including self-closing). |
| role_delimiter_injection | Injecting role delimiters like `[system]:`. |
| bracketed_role_spoofing | Fake bracketed role labels (e.g. `[System]`, `[Assistant]`). |
| system_prefix_spoofing | Lines starting with `System:` to impersonate system messages (multiline). |

### Control Token Injection

| Pattern | Detects |
|---|---|
| control_token_injection | ChatML / Llama 3 / generic pipe-delimited control tokens. |
| deepseek_control_token_injection | DeepSeek fullwidth-pipe (`｜`) control tokens. |

## Custom patterns

contentFilters holds up to 20 entries of { pattern, action, label }. pattern is
a JavaScript regular expression source of 1 to 2000 characters, compiled with
the g and u flags. Matching is case-sensitive: write [Aa]cme to match either
case. label is optional, up to 40 letters, digits, spaces, _ or -; the
redaction placeholder is [label], or [REDACTED] without a label. Events list
the label, or "pattern n" for an unlabeled entry.

A pattern is rejected with 400 when it uses lookahead, lookbehind or
backreferences (\1, \k<name>); nests a quantifier inside a quantified group
such as (a+)+; matches the empty string; or runs too slowly on plain text in
the save-time probe. The message reads "Pattern n: <rule>" and never repeats
the pattern. Custom patterns scan every message role.

## What is scanned

Only these strings are read and rewritten, on /chat/completions, /responses,
/messages and /messages/count_tokens. Anything else at these paths, and every
other field, is left alone. The scan runs before token measurement, the
balance hold, the request trace and the upstream send, so all of them see the
redacted text.

| Route | Path | Role for scan scope |
|---|---|---|
| /chat/completions | messages[i].content when it is a string | messages[i].role |
|  | messages[i].content[j].text when content[j].type is text | messages[i].role |
|  | messages[i].tool_calls[j].function.arguments when it is a string | assistant |
| /responses | input when it is a string | user |
|  | input[i].content when it is a string | input[i].role, else user |
|  | input[i].content[j].text when the type is input_text or output_text | input[i].role, else user |
|  | input[i].arguments when input[i].type is function_call | assistant |
|  | instructions when it is a string | system |
| /messages, /messages/count_tokens | system when it is a string; system[i].text when system[i].type is text | system |
|  | messages[i].content when it is a string | messages[i].role |
|  | messages[i].content[j].text when the type is text | messages[i].role |
|  | messages[i].content[j].content when the type is tool_result and it is a string | user |

Input only: model output is never scanned. Embeddings, rerank and image
requests are not content-scanned; they still get the merged model and
provider lists and the budget.

## When a request is blocked

Block filters run first; any hit refuses the request with 403
guardrail_blocked before anything is measured, held or sent, and nothing is
charged. The message names the detector or pattern, never the text it
matched, and is capped at 200 characters. The hit appears in the hits list
on https://minirouter.sh/dashboard/guardrails and on GET /guardrails/events;
hits are kept for 30 days.

On /chat/completions and /responses:

    {
      "error": {
        "message": "A guardrail on this API key blocked the request: prompt injection patterns detected (ignore_previous_instructions). Nothing was sent to a provider and nothing was charged. Manage guardrails at https://minirouter.sh/dashboard/guardrails Docs: https://minirouter.sh/docs/guardrails",
        "type": "invalid_request_error",
        "code": "guardrail_blocked",
        "request_id": "req_01J8Z6Q0GUARDRAIL"
      }
    }

On /messages and /messages/count_tokens:

    {
      "type": "error",
      "error": {
        "type": "permission_error",
        "message": "A guardrail on this API key blocked the request: prompt injection patterns detected (ignore_previous_instructions). Nothing was sent to a provider and nothing was charged. Manage guardrails at https://minirouter.sh/dashboard/guardrails Docs: https://minirouter.sh/docs/guardrails"
      },
      "request_id": "req_01J8Z6Q0GUARDRAIL"
    }

Error reference: https://minirouter.sh/docs/errors#guardrail_blocked

## Headers

x-minirouter-guardrail is set when a guardrail acted without blocking. A
block is a 403, never a header; a request nothing matched carries no header unless the scan failed (`skipped`).

- redacted: Text was replaced before the request was measured, held against your balance or sent upstream.
- flagged: A flag filter matched. The request went through unchanged and the hit was recorded.
- skipped: The scanner failed on this request. It was served unscanned and the failure was logged.

## API

Base URL https://api.minirouter.sh, the same as /keys. Reads accept an API
key or a signed-in session; writes and the test endpoint need a signed-in
session. Writes are limited to 30 per minute per account.

| Endpoint | Principal | Behaviour |
|---|---|---|
| GET /guardrails | key or session | Every guardrail on the account plus limits { maxGuardrails: 20, maxCustomPatterns: 20 }. |
| POST /guardrails | session | Body: name, description and any config field; omitted fields start empty. 201 with the guardrail. 409 for a duplicate name or a 21st guardrail. |
| GET /guardrails/:id | key or session | One guardrail. 404 when the id is not on this account. |
| PATCH /guardrails/:id | session | Body: expectedRevision plus the fields to change. Arrays replace, never merge. 409 policy_revision_conflict when the revision is stale. |
| DELETE /guardrails/:id | session | Hard delete. Keys that used it are unassigned. Returns { deleted: true, unassignedKeys }. |
| POST /guardrails/test | session | Body: text (up to 20,000 characters) plus exactly one of guardrailId or config. Scans text as a user turn and returns { action, redacted, hits }. Nothing is stored or logged. |
| GET /guardrails/events | key or session | Hits, newest first. Query: limit (1 to 200, default 50), cursor (the nextCursor of the previous page), guardrailId. |

Fields on POST and PATCH:

- name: Unique on the account, up to 80 characters.
- description: Optional, up to 400 characters.
- limitNano: Budget in nano-USD as a decimal string, or null. Sent together with resetInterval.
- resetInterval: daily, weekly or monthly, or null.
- allowedModels: Catalog ids or author/* families. Empty means every model. Up to 317 entries.
- ignoredModels: Same grammar. These models are refused even when allowed elsewhere.
- allowedProviders: Provider ids. Empty means every provider.
- ignoredProviders: Provider ids to remove.
- contentFilterBuiltins: Array of { slug, action }. One entry per detector; a detector not listed is off.
- contentFilters: Array of { pattern, action, label }. Up to 20 patterns.
- injectionScanScope: user_only (default) or all_messages.

Every guardrail is returned with id, revision (send it back as
expectedRevision), keyCount, createdAt and updatedAt. Unknown fields are a
400 naming the field.

Events carry id, guardrailId, guardrailName (null once deleted), apiKeyId,
keyPrefix, requestId, route, detector (a preset slug or custom), action,
patterns (pattern names, format ids or labels) and matchCount. Pages return
nextCursor; pass it as cursor for the next page, null means the end.
