Guide
Set a ceiling on agent spend.
Contain a runaway loop with prepaid credits, per-key spend caps, and a separate key for each agent.
Ceiling one — the balance itself
Credits are prepaid. When they are gone the requests stop, and there is no invoice arriving afterwards for the part you did not notice.
Prepaid credits bound your spend by the amount you deposit. A runaway agent cannot keep billing your card.
The corollary is worth stating in the same breath: there are no withdrawals. Credits buy inference and never convert back to coin. Keeping a small balance is the containment mechanism, not a cash position.
Ceiling two — per-key caps
Four limits are enforced at the edge, per key. Open reservations count against them the moment a request starts, so a burst of concurrent calls cannot slip past a cap that a serial one would have hit.
- FieldrpmLimitCapsrequests per minute
- FieldtpmLimitCapsestimated input and output tokens per minute
- FielddailyLimitNanoCapsdaily spend cap in nano-USD
- FieldmonthlyLimitNanoCapsmonthly spend cap in nano-USD
Defaults are lower for accountless keys and every one of them can be raised per key. When a request is refused the 429 carries Retry-After and names which limit was reached and where to raise it — see /docs/rate-limits.
What to actually set the cap to
Cost one turn, multiply by the turns you expect in a day, then set the cap somewhere above that and below what you would mind losing. The example below uses current catalog rates; replace its token counts and turn frequency with your own workload.
- Example input per turn
- 4K tokens
- Example output per turn
- 250 tokens
- Example runaway
- 480 turns
System prompt, tool schema and recent history, re-sent each turn.
A short reply or a tool call, not an essay.
One turn a minute for 8 hours — an overnight loop nobody was awake for.
| Item | USD |
|---|---|
DeepSeek V4 Flash 0731$0.17 · 31% of chart maximum
One turn costs $0.0004 on DeepSeek V4 Flash 0731 and $0.0012 on GPT 5.6 Luna. The useful reading of that spread is not which model is cheaper — it is that the same cap protects you very differently depending on the model behind it, so set the cap after you have chosen the model, not before.
Round up, not down. A cap that trips during normal operation trains you to raise it without thinking, which is how caps stop working. Set it at several times your expected day so that hitting it is genuinely information.
Ceiling three — one key per agent
The only one of the three that is a habit rather than a mechanism, and the one that does the most work.
Give each agent its own key and limit. If one hits its cap, the others keep working and you know which agent to inspect.
What hitting a ceiling looks like
Your agent's errors get read inside a chat app where we have no UI, so each one has to be self-contained.
curl https://api.minirouter.sh/v1/chat/completions \
-H "Authorization: Bearer $MINIROUTER_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"meta/llama-3.3-70b","messages":[{"role":"user","content":"ping"}]}' \
-D - -o /dev/null
# every response carries the balance, so an agent can warn you
# before it runs dry rather than after:
# x-minirouter-balance-usd: 4.81
# x-minirouter-cost-usd: 0.0004
# x-minirouter-provider: together
# x-minirouter-ttft-ms: 240- Before any tokens arrive — a 402 naming the shortfall, if the balance is below the estimated cost of the request.
- Mid-stream — the stream ends with a terminal error stating the exact charge for tokens already delivered, then closes cleanly. Never a silent truncation you have to diagnose.
- At a rate cap — a 429 with
Retry-After, naming the limit reached and where to raise it. - Every one of them carries a URL, because inside someone else’s chat pane the URL is the only navigation there is. /docs/errors
Next
- Personal agentsClient guides for OpenClaw, Hermes Agent and OpenHands.
- Rate limits referenceEvery field, the reservation rule, and the 429 shape.
- Streaming and abortsThe cost trailer, and what you are charged when you press Stop.
- BillingHow credits are priced, and what happens at zero.
- Cost per messageThe same arithmetic for a very different workload shape.
- The full catalogCurrent rates for every model, to redo this sum with your own.