Integrations
Hermes Agent
Run Hermes Agent on any MiniRouter model through its custom endpoint.
Hermes Agent is Nous Research's always-on personal agent. Add MiniRouter as its custom endpoint to run it on any model in the catalog, move side tasks to cheaper models, and cap unattended spend with rate limits.
1. Install Hermes
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bashiex (irm https://hermes-agent.nousresearch.com/install.ps1)2. Save your key
Create a key, then add it to Hermes's env file.
MINIROUTER_KEY=mr-live-YOUR-KEY-HERE3. Point Hermes at MiniRouter
hermes model- hermes model
- Custom endpoint (self-hosted / VLLM / etc.)
- model.
base_url https://api.minirouter.sh/v1Paste it complete, including /v1.- model.
provider custom- model.
default zai/glm-5.3-flash- model.
key_env MINIROUTER_KEY- MINIROUTER_KEY
mr-live-… (your key)In ~/.hermes/.env, not config.yaml.
model:
provider: custom
base_url: https://api.minirouter.sh/v1
default: zai/glm-5.3-flash
key_env: MINIROUTER_KEY4. Start and verify
hermesSend one message, then open Activity. The request lists its model, tokens and exact cost. If nothing arrives, run hermes doctor.
Switch models
- /
model custom: openai/ gpt-5. 6-luna - This session.
- /
model custom: openai/ gpt-5. 6-sol --once - The next turn only.
- /
model custom: openai/ gpt-5. 6-luna --global - This session and new ones.
- hermes chat -m openai/
gpt-5. 6-luna - One run.
- model.
default - New sessions, in
~/.hermes/config.yaml.
Run from scripts
-m overrides the model for that run only.
hermes -z "Summarize today's inbox" \
-m zai/glm-5.3-flash \
--usage-file usage.json- hermes -z
- The final answer only.
- hermes chat --oneshot -q
- The answer and tool output, then exits.
Move side tasks and add fallbacks
Compression, vision and titles run on your main model unless you move them. A fallback takes over when the main model fails.
auxiliary:
compression:
base_url: https://api.minirouter.sh/v1
model: deepseek/deepseek-v4-flash-0731
api_key: ${MINIROUTER_KEY}
fallback_providers:
- provider: custom
model: openai/gpt-5.6-luna
base_url: https://api.minirouter.sh/v1
key_env: MINIROUTER_KEYCompare rates for side-task models on Models, or read Model fallbacks for server-side fallbacks.
Keep spend in check
Restrict models or filter content with Guardrails.
Recommended modelsBrowse every model in the catalog.
zai/glm-5.3-flash- Strong agentic performance at a low input rate for repeated tool schemas.
openai/gpt-5.6-luna- Low-cost, high-intelligence alternative for routine agent turns.
deepseek/deepseek-v4-flash-0731- Lowest input rate among the high-scoring long-context candidates.
tencent/hy3- Lower-capability fallback with inexpensive input for high-volume loops.
Troubleshooting
Provider loads but every model call 404s
Paste https://api.minirouter.sh/v1 complete, including the /v1. These clients treat the base URL as the OpenAI-compatible root and append /chat/completions themselves, so trimming the /v1 leaves them calling a path that does not exist.
Error: 400 invalid_request.
Credits gone overnight with nothing to show
An agent in a retry loop spends at machine speed. Set dailyLimitNano on the key the agent uses — one key per agent, so a runaway is contained to that key rather than the whole balance.
Error: 402 insufficient_credits.
401 on every request
Wrong or rotated key. Keys start with mr-live-; after a rotation the old key keeps working for 24 hours, then dies.
Error: 401 invalid_api_key.
Response stops mid-sentence with a credits message
Balance hit $0 mid-stream. The stream ends with a terminal error naming the exact charge for tokens delivered — never a silent close.
Error: 402 insufficient_credits.
Every error code, with its fix: Errors.